[Paper Review] Exploiting Smoothness in Statistical Learning, Sequential Prediction, and Stochastic Optimization
This dissertation introduces a unified framework for exploiting smoothness in statistical learning, online learning, and stochastic optimization. By leveraging smoothness through novel algorithms—including mixed optimization and projection-free methods—it achieves exponential improvements in sample complexity, condition-number-independent convergence, and regret bounds tied to gradual variation, while maintaining statistical consistency and computational efficiency across settings.
In the last several years, the intimate connection between convex optimization and learning problems, in both statistical and sequential frameworks, has shifted the focus of algorithmic machine learning to examine this interplay. In particular, on one hand, this intertwinement brings forward new challenges in reassessment of the performance of learning algorithms including generalization and regret bounds under the assumptions imposed by convexity such as analytical properties of loss functions (e.g., Lipschitzness, strong convexity, and smoothness). On the other hand, emergence of datasets of an unprecedented size, demands the development of novel and more efficient optimization algorithms to tackle large-scale learning problems. The overarching goal of this thesis is to reassess the smoothness of loss functions in statistical learning, sequential prediction/online learning, and stochastic optimization and explicate its consequences. In particular we examine how smoothness of loss function could be beneficial or detrimental in these settings in terms of sample complexity, statistical consistency, regret analysis, and convergence rate, and investigate how smoothness can be leveraged to devise more efficient learning algorithms.
Motivation & Objective
- To reassess the role of smoothness in statistical learning, online learning, and stochastic optimization.
- To develop algorithms that exploit smoothness to improve convergence rates and sample complexity.
- To analyze the trade-offs between smoothness and statistical consistency in surrogate loss functions.
- To design efficient optimization methods that reduce reliance on full gradients while maintaining fast convergence.
- To enable projection-free and constraint-aware algorithms for large-scale learning problems.
Proposed method
- Introduces mixed optimization, which interpolates between stochastic and full gradient methods by using infrequent full gradients to reduce stochastic gradient variance.
- Proposes a new performance measure—gradual variation—defined as the sum of distances between consecutive loss functions in online learning.
- Designs projection-free algorithms that require only one projection at the final iteration, reducing computational cost in stochastic optimization.
- Develops online convex optimization with soft constraints, allowing sub-linear regret and constraint violation bounds.
- Uses a constructive proof based on a properly designed stochastic optimization algorithm to establish exponential sample complexity improvement.
- Provides a unified analysis of optimization error, generalization error, and binary excess risk to identify conditions under which smoothness is beneficial.
Experimental results
Research questions
- RQ1How does smoothness of the loss function affect sample complexity in statistical learning when the target risk is known?
- RQ2In what conditions does smoothness of surrogate loss functions improve or degrade statistical consistency and binary excess risk?
- RQ3Can smoothness be exploited to achieve faster convergence in stochastic optimization without full gradient access?
- RQ4How can gradual variation in loss functions be leveraged to improve regret bounds in online learning?
- RQ5Can projection-free and constraint-aware algorithms be designed to maintain sub-linear regret while minimizing computational overhead?
Key findings
- Under smoothness and strong convexity, sample complexity improves exponentially when the target risk is known, via a constructive stochastic optimization algorithm.
- Smoothness of surrogate loss functions can deteriorate binary excess risk, highlighting a trade-off between computational efficiency and statistical consistency.
- The proposed mixed optimization method achieves condition-number-independent convergence rates in deterministic optimization and faster convergence in stochastic settings.
- Regret bounds in online convex optimization are improved by leveraging gradual variation, enabling performance gains in non-adversarial, smoothly evolving environments.
- Projection-free algorithms require only one projection at the final iteration, significantly reducing computational cost in stochastic optimization.
- Soft-constrained online learning achieves sub-linear regret and sub-linear constraint violation, enabling efficient learning under long-term feasibility constraints.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.