Skip to main content
QUICK REVIEW

[Paper Review] Limitations of adversarial robustness: strong No Free Lunch Theorem.

Elvis Dohmatob|arXiv (Cornell University)|Oct 8, 2018
Adversarial Robustness in Machine Learning14 references15 citations
TL;DR

This paper establishes a Strong No Free Lunch Theorem for adversarial robustness, proving that any classifier on data satisfying a $W_2$ Talagrand transportation-cost inequality—such as log-concave or uniform distributions on positively curved manifolds—can be adversarially fooled with high probability when perturbations exceed natural noise levels. The result generalizes prior impossibility results and is validated on MNIST and simulated data.

ABSTRACT

This manuscript presents some new impossibility results on adversarial robustness in machine learning, a very important yet largely open problem. We show that if conditioned on a class label the data distribution satisfies the $W_2$ Talagrand transportation-cost inequality (for example, this condition is satisfied if the conditional distribution has density which is log-concave; is the uniform measure on a compact Riemannian manifold with positive Ricci curvature, any classifier can be adversarially fooled with high probability once the perturbations are slightly greater than the natural noise level in the problem. We call this result The Strong No Free Lunch Theorem as some recent results (Tsipras et al. 2018, Fawzi et al. 2018, etc.) on the subject can be immediately recovered as very particular cases. Our theoretical bounds are demonstrated on both simulated and real data (MNIST). We conclude the manuscript with some speculation on possible future research directions.

Motivation & Objective

  • To identify fundamental limitations in achieving adversarial robustness across natural data distributions.
  • To formalize conditions under which adversarial examples are inevitable, even with optimal classifiers.
  • To generalize prior impossibility results into a broader theoretical framework using transportation-cost inequalities.
  • To validate theoretical bounds empirically on MNIST and synthetic data.
  • To guide future research by identifying intrinsic barriers to robust learning.

Proposed method

  • Theoretical analysis is grounded in the $W_2$ Talagrand transportation-cost inequality, which characterizes the concentration of conditional data distributions.
  • The paper assumes that for each class label, the data distribution satisfies the $W_2$ Talagrand inequality, a condition met by log-concave densities and uniform measures on compact Riemannian manifolds with positive Ricci curvature.
  • It proves that under this condition, any classifier becomes vulnerable to adversarial attacks when perturbations exceed the natural noise level.
  • The proof leverages geometric and probabilistic tools from optimal transport and concentration of measure.
  • Theoretical bounds are derived and empirically tested on MNIST and simulated data with log-concave and uniform distributions.
  • The framework unifies and extends prior results (e.g., Tsipras et al. 2018, Fawzi et al. 2018) as special cases.

Experimental results

Research questions

  • RQ1Under what distributional assumptions is adversarial robustness fundamentally impossible?
  • RQ2Can the $W_2$ Talagrand transportation-cost inequality be used to derive universal lower bounds on adversarial vulnerability?
  • RQ3To what extent do existing robustness results depend on specific data assumptions that are not generalizable?
  • RQ4How do natural noise levels relate to the threshold for adversarial vulnerability in realistic data distributions?
  • RQ5Can the proposed framework recover and generalize prior impossibility results in adversarial robustness?

Key findings

  • Any classifier trained on data satisfying the $W_2$ Talagrand transportation-cost inequality becomes vulnerable to adversarial attacks when perturbations exceed the natural noise level.
  • The result applies to a broad class of distributions, including log-concave densities and uniform measures on compact Riemannian manifolds with positive Ricci curvature.
  • The theoretical bounds on adversarial vulnerability are demonstrated to hold on both simulated data and MNIST, confirming the predictions empirically.
  • The framework recovers and generalizes prior impossibility results such as those by Tsipras et al. (2018) and Fawzi et al. (2018) as special cases.
  • The findings suggest that adversarial robustness is not achievable in principle under mild and natural assumptions on data distributions.
  • The study identifies a fundamental trade-off between natural accuracy and adversarial robustness under these distributional conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.