[Paper Review] Minimax Statistical Learning with Wasserstein Distances
This paper introduces a minimax statistical learning framework using Wasserstein balls as ambiguity sets, derives data-dependent generalization bounds, and applies the approach to transport-based domain adaptation without explicitly estimating transport maps.
As opposed to standard empirical risk minimization (ERM), distributionally robust optimization aims to minimize the worst-case risk over a larger ambiguity set containing the original empirical distribution of the training data. In this work, we describe a minimax framework for statistical learning with ambiguity sets given by balls in Wasserstein space. In particular, we prove generalization bounds that involve the covering number properties of the original ERM problem. As an illustrative example, we provide generalization guarantees for transport-based domain adaptation problems where the Wasserstein distance between the source and target domain distributions can be reliably estimated from unlabeled samples.
Motivation & Objective
- Motivate distributional robustness by guarding against domain drift in learning.
- Introduce local minimax risk with Wasserstein ambiguity sets centered at the empirical distribution.
- Derive data-dependent generalization bounds for minimax ERM and excess risk under smoothness assumptions.
- Provide a domain adaptation bound in the optimal transport framework that relies on unlabeled data.
- Connect the Wasserstein-based approach to existing domain adaptation and robust optimization literature.
Proposed method
- Define the p-Wasserstein ball around the empirical distribution as the ambiguity set A(Pn).
- Formulate local minimax risk Rrho,p(P,f) = supQ in Aw of P R(Q,f) and the corresponding minimax ERM objective.
- Utilize Kantorovich duality to relate Rrho,p(Q,f) to an optimization over lambda and a surrogate function phi_{lambda,f}.
- Derive data-dependent generalization bounds for Rrho,p(P,f) in terms of covering numbers (entropy integral) of F (Theorem 1).
- Obtain excess risk bounds under uniform smoothness (Theorem 2) and under a weaker assumption that at least one smooth f0 exists (Theorem 3).
- Present a domain adaptation bound (Theorem 4) within the Courty et al. transport framework using unlabeled target data and estimated Wasserstein distance.
Experimental results
Research questions
- RQ1How does local minimax risk under Wasserstein ambiguity relate to traditional statistical risk?
- RQ2Can empirical risk minimization under Wasserstein-based ambiguity achieve near-optimal local minimax risk?
- RQ3What are data-dependent and parameter-free (or minimally parameterized) generalization/excess risk bounds under Wasserstein ambiguity?
- RQ4How can Wasserstein-based distributional robustness provide guarantees for domain adaptation without explicit transport-map estimation?
Key findings
- The local minimax ERM performs close to the optimal local minimax risk, with generalization bounds that adapt to the ambiguity radius and the complexity of the hypothesis class.
- When rho = 0, the bounds recover ordinary ERM rates, showing consistency with standard learning in the absence of ambiguity.
- Under uniform smoothness, an explicit excess risk bound scales with the entropy integral and the ambiguity radius, clarifying the trade-off between robustness and learnability.
- Even with weaker assumptions, there exist meaningful excess risk bounds that still exhibit favorable dependence on rho and sample size.
- In domain adaptation, the framework yields a target-domain generalization bound expressed via the estimated Wasserstein distance between source and target distributions, without requiring direct transport-map estimation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.