Entropy Regularization for Control and Estimation
Speaker
About this event
```Control and estimation for stochastic dynamical systems concern the synthesis of input policies and the reconstruction of state information from measurements. Entropy provides a quantitative description of the randomness present in system trajectories, disturbance models, and belief distributions. The main contribution of this thesis is the development of entropy-regularized methods for stochastic control and estimation. This involves topics such as dynamic programming, formal abstractions, and recursive Bayesian filtering. The central contribution concerns entropy-regularized control of continuous-state stochastic systems through finite abstractions. Existing abstraction methods enable formal controller synthesis for objectives such as cumulative costs and temporal-logic specifications. Entropy-based trajectory objectives do not transfer through these abstractions directly, as discretization changes the entropy of the induced trajectory distribution. This thesis derives bounds relating the Kullback-Leibler (KL) divergence to uniform of a continuous trajectory distribution to that of its finite discretization. These bounds enable formal entropy-aware controller synthesis for continuous-state systems, trading cumulative cost against trajectory predictability while retaining guarantees for the original system. A second contribution concerns robust stochastic control. It generalizes KL-regularized robust-control formulations by allowing adversarial disturbance distributions to be regularized through both cross entropy with respect to an empirical model and the entropy of the adversary itself. The resulting dynamic-programming recursion gives rise to the minsoftmax algorithm and places minimax, stochastic, KL-regularized, and H-infinity type viewpoints in one parameterized formulation. A third contribution concerns Bayesian filtering. The tempered Bayes filter modifies the recursive Bayesian update by tempering the distributions that define the posterior. This yields a computationally efficient modification of the Bayes filter that can improve predictive performance under model mismatch. Specializing the construction to the linear Gaussian setting yields the tempered Kalman filter. Two further papers are included as supporting material. The entropy-regularized interval Markov decision process work supports the continuous-state abstraction chapter. The optimal stopping work is included as a thesis appendix outside the main entropy-regularization arc.```
Host