[Paper Review] Improper Learning for Non-Stochastic Control
This paper introduces a novel controller parametrization using denoised observations and applies online gradient descent to achieve sublinear regret in non-stochastic control with partial observations. It establishes the first √T regret bound for known systems and T^{2/3} for unknown systems, along with optimal poly(log T) regret under semi-adversarial noise, competing with all stabilizing linear controllers.
We consider the problem of controlling a possibly unknown linear dynamical system with adversarial perturbations, adversarially chosen convex loss functions, and partially observed states, known as non-stochastic control. We introduce a controller parametrization based on the denoised observations, and prove that applying online gradient descent to this parametrization yields a new controller which attains sublinear regret vs. a large class of closed-loop policies. In the fully-adversarial setting, our controller attains an optimal regret bound of $\sqrt{T}$-when the system is known, and, when combined with an initial stage of least-squares estimation, $T^{2/3}$ when the system is unknown; both yield the first sublinear regret for the partially observed setting. Our bounds are the first in the non-stochastic control setting that compete with \emph{all} stabilizing linear dynamical controllers, not just state feedback. Moreover, in the presence of semi-adversarial noise containing both stochastic and adversarial components, our controller attains the optimal regret bounds of $\mathrm{poly}(\log T)$ when the system is known, and $\sqrt{T}$ when unknown. To our knowledge, this gives the first end-to-end $\sqrt{T}$ regret for online Linear Quadratic Gaussian controller, and applies in a more general setting with adversarial losses and semi-adversarial noise.
Motivation & Objective
- To address non-stochastic control with adversarial dynamics, losses, and partial state observation, where traditional optimal controllers cannot be pre-computed.
- To develop a regret-minimizing controller that competes with the broad class of stabilizing linear dynamical controllers, not just static feedback policies.
- To unify and improve prior results in non-stochastic control, extending to settings with both stochastic and adversarial noise.
- To achieve tight regret bounds in the classical Linear Quadratic Gaussian (LQG) setting with unknown systems, where prior work lacked √T regret guarantees.
Proposed method
- Proposes a controller parametrization based on the Youla parametrization, re-expressing the system in terms of 'Nature’s y’s'—observations under zero control.
- Introduces the Disturbance Response Control (DRC) framework, a convex parametrization of stabilizing controllers derived from denoised observations.
- Applies online gradient descent to the DRC parametrization to form the Gradient Response Controller via Gradient Descent (DRC-GD).
- Uses a two-stage approach: initial least-squares estimation for unknown systems, followed by online optimization over the DRC space.
- Employs martingale concentration and sub-Gaussian tail bounds to derive high-probability regret bounds under adversarial and semi-adversarial noise.
- Derives regret bounds via recursive decomposition of loss differences and careful control of gradient and state estimation errors.
Experimental results
Research questions
- RQ1Can sublinear regret be achieved in non-stochastic control with partial observations when the system dynamics are unknown?
- RQ2Does the proposed controller achieve optimal regret rates in both fully adversarial and semi-adversarial noise settings?
- RQ3Can the regret bound compete with the entire class of stabilizing linear dynamical controllers, not just state-feedback policies?
- RQ4Is the √T regret bound tight for the known system case in the partially observed setting?
- RQ5Can the framework achieve poly(log T) regret under semi-adversarial noise when the system is known?
Key findings
- The DRC-GD controller achieves Õ(√T) regret for known systems under full adversarial conditions, which is tight and the first such bound in the partially observed setting.
- For unknown systems, the method achieves Õ(T^{2/3}) regret, the first sublinear regret bound in this setting with partial observation.
- In the presence of semi-adversarial noise (stochastic + adversarial), the controller attains poly(log T) regret when the system is known, and √T when unknown—both optimal.
- The regret bounds hold against the full class of stabilizing linear dynamical controllers, a significantly richer class than previously considered static feedback policies.
- The framework provides the first end-to-end √T regret bound for online Linear Quadratic Gaussian (LQG) control with an unknown system.
- The results extend to general convex loss functions and adversarial perturbations, demonstrating robustness and generality.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.