[Paper Review] A Stochastic View of Optimal Regret through Minimax Duality
This paper establishes a stochastic interpretation of optimal regret in online convex optimization using von Neumann's minimax theorem, showing that the minimax regret equals the maximum gap in Jensen’s inequality for a concave functional over probability distributions of adversarial data sequences. It derives tight upper and lower bounds on regret through geometric and duality-based analysis, revealing connections to statistical learning and providing explicit optimal adversary strategies.
We study the regret of optimal strategies for online convex optimization games. Using von Neumann's minimax theorem, we show that the optimal regret in this adversarial setting is closely related to the behavior of the empirical minimization algorithm in a stochastic process setting: it is equal to the maximum, over joint distributions of the adversary's action sequence, of the difference between a sum of minimal expected losses and the minimal empirical loss. We show that the optimal regret has a natural geometric interpretation, since it can be viewed as the gap in Jensen's inequality for a concave functional--the minimizer over the player's actions of expected loss--defined on a set of probability distributions. We use this expression to obtain upper and lower bounds on the regret of an optimal strategy for a variety of online learning problems. Our method provides upper bounds without the need to construct a learning algorithm; the lower bounds provide explicit optimal strategies for the adversary.
Motivation & Objective
- To bridge adversarial online learning and statistical learning by revealing a deep duality between optimal regret and empirical minimization in stochastic processes.
- To characterize the minimax regret in online convex optimization as a geometric gap in Jensen’s inequality for a concave functional over distributions.
- To derive upper and lower bounds on optimal regret without constructing learning algorithms, using duality and support function analysis.
- To identify explicit optimal adversary strategies that achieve the derived lower bounds on regret.
- To explore invariance of regret under linear transformations of the loss class, linking isomorphic learning problems.
Proposed method
- Applying von Neumann’s minimax theorem to reframe the minimax regret as an expectation gap between conditional minimal losses and global minimal empirical loss.
- Defining the regret as the supremum over joint distributions of the data sequence of the difference between sum of conditional minima and the global empirical minimum.
- Using the support function of the convex hull of negative loss functions to characterize the regret via convex duality.
- Applying backward induction on a constructed sequence of conditional expectations to prove the exact regret expression for a specific distribution.
- Leveraging known asymptotic expansions (e.g., ∑cₜ = log T − log log T + o(1)) to derive the exact asymptotic regret rate.
- Using linear transformations of the loss class to analyze invariance of regret under invertible mappings, preserving the √T rate if exposed faces are preserved.
Experimental results
Research questions
- RQ1How is the optimal regret in online convex optimization related to the behavior of empirical minimization in a stochastic setting?
- RQ2Can the minimax regret be expressed as a geometric gap in Jensen’s inequality for a concave functional over probability distributions?
- RQ3What are the tight upper and lower bounds on optimal regret for online learning problems, and how can they be derived without constructing algorithms?
- RQ4What explicit strategies does the adversary use to achieve the worst-case regret, and how are they characterized?
- RQ5How do linear transformations of the loss class affect the minimax regret, and when is the regret invariant under such transformations?
Key findings
- The minimax regret is exactly equal to the supremum over all joint distributions of the data sequence of the difference between the sum of conditional minimal expected losses and the minimal empirical loss.
- The optimal regret exhibits a precise asymptotic form: log T − log log T + o(1), derived via backward induction on a constructed sequence of conditional expectations.
- The regret can be interpreted as the gap in Jensen’s inequality for the concave functional defined as the infimum of expected loss over the player’s action set.
- The method provides lower bounds via explicit optimal adversary strategies, which are derived from the optimal distribution over data sequences.
- Upper bounds on regret are obtained without constructing learning algorithms, relying solely on duality and convex geometry.
- Linear transformations of the loss class preserve the √T regret rate if they do not rotate away non-singular exposed faces from the origin, indicating isomorphism of certain learning problems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.