[Paper Review] Stochastic Hamiltonian Gradient Methods for Smooth Games
This paper introduces stochastic Hamiltonian gradient descent (SHGD) with a novel unbiased estimator for stochastic smooth games, providing the first global non-asymptotic last-iterate convergence guarantees for unconstrained bilinear and 'sufficiently bilinear' games without strong monotonicity. It further proposes L-SVRHG, the first stochastic variance-reduced Hamiltonian method, achieving linear convergence under mild conditions.
The success of adversarial formulations in machine learning has brought renewed motivation for smooth games. In this work, we focus on the class of stochastic Hamiltonian methods and provide the first convergence guarantees for certain classes of stochastic smooth games. We propose a novel unbiased estimator for the stochastic Hamiltonian gradient descent (SHGD) and highlight its benefits. Using tools from the optimization literature we show that SHGD converges linearly to the neighbourhood of a stationary point. To guarantee convergence to the exact solution, we analyze SHGD with a decreasing step-size and we also present the first stochastic variance reduced Hamiltonian method. Our results provide the first global non-asymptotic last-iterate convergence guarantees for the class of stochastic unconstrained bilinear games and for the more general class of stochastic games that satisfy a "sufficiently bilinear" condition, notably including some non-convex non-concave problems. We supplement our analysis with experiments on stochastic bilinear and sufficiently bilinear games, where our theory is shown to be tight, and on simple adversarial machine learning formulations.
Motivation & Objective
- Address the lack of convergence guarantees for stochastic Hamiltonian methods in smooth games, especially without strong monotonicity.
- Provide theoretical convergence analysis for SHGD and L-SVRHG in stochastic settings, focusing on last-iterate convergence.
- Overcome limitations of biased estimators in existing SHGD variants by introducing an unbiased gradient estimator for the Hamiltonian function.
- Extend convergence theory to non-compact domains and non-convex non-concave problems via the 'sufficiently bilinear' condition.
- Demonstrate tightness of theory through experiments on bilinear, sufficiently bilinear, and adversarial learning games.
Proposed method
- Propose a novel unbiased estimator for the stochastic Hamiltonian gradient, enabling rigorous convergence analysis.
- Apply standard optimization tools (e.g., Lyapunov functions) to prove linear convergence of SHGD with constant step-size to a neighborhood of the stationary point.
- Introduce a decreasing step-size schedule to achieve sub-linear convergence to the exact stationary point.
- Design L-SVRHG, the first stochastic variance-reduced Hamiltonian method, using a control variate technique to reduce gradient variance.
- Formulate the stochastic game as a finite-sum problem: $ g(x_1,x_2) = \frac{1}{n}\sum_{i=1}^n g_i(x_1,x_2) $, with each $ g_i $ smooth.
- Analyze convergence under the 'sufficiently bilinear' condition, which generalizes bilinear games and includes some non-monotone problems.
Experimental results
Research questions
- RQ1Can stochastic Hamiltonian gradient descent achieve last-iterate convergence without strong monotonicity assumptions?
- RQ2What is the convergence rate of SHGD with constant and decreasing step-sizes in stochastic smooth games?
- RQ3Can variance reduction be effectively applied to Hamiltonian methods in stochastic games?
- RQ4How does the proposed unbiased estimator improve convergence guarantees compared to biased variants?
- RQ5Does the theory hold tightly in practice, especially for adversarial learning and bilinear games?
Key findings
- SHGD with constant step-size converges linearly to a neighborhood of the stationary point under the 'sufficiently bilinear' condition.
- SHGD with a decreasing step-size achieves sub-linear convergence to the exact stationary point, ensuring global convergence.
- L-SVRHG achieves linear convergence with a lower iteration cost than standard SVRG, requiring $ 4 + p \cdot 2n $ gradient computations per iteration.
- Experiments show that the theoretical convergence rates are tight, with L-SVRHG being the fastest method in bilinear and sufficiently bilinear games.
- In the interpolation setting (where $ \xi_i(x^*) = 0 $), all methods converge linearly, and surprisingly, biased SHGD outperforms others due to signal consistency.
- The proposed unbiased estimator enables convergence guarantees where previous biased variants fail to ensure last-iterate convergence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.