Skip to main content
QUICK REVIEW

[Paper Review] Symmetry & critical points for a model shallow neural network

Yossi Arjevani, Michael Field|arXiv (Cornell University)|Mar 23, 2020
Neural Networks and Applications25 references4 citations
TL;DR

This paper analyzes critical points in two-layer ReLU neural networks with k hidden neurons using symmetry and group theory, expressing them as power series in $k^{-1/2}$. It reveals that spurious minima vary significantly: some have loss decaying as $k^{-1}$, while others converge to a positive constant, demonstrating that not all spurious minima are equivalent in optimization behavior.

ABSTRACT

We consider the optimization problem associated with fitting two-layer ReLU networks with $k$ hidden neurons, where labels are assumed to be generated by a (teacher) neural network. We leverage the rich symmetry exhibited by such models to identify various families of critical points and express them as power series in $k^{-\frac{1}{2}}$. These expressions are then used to derive estimates for several related quantities which imply that not all spurious minima are alike. In particular, we show that while the loss function at certain types of spurious minima decays to zero like $k^{-1}$, in other cases the loss converges to a strictly positive constant. The methods used depend on symmetry, the geometry of group actions, bifurcation, and Artin's implicit function theorem.

Motivation & Objective

  • To understand the structure and diversity of critical points in two-layer ReLU networks trained on data generated by a teacher network.
  • To classify critical points based on their geometric and symmetric properties under group actions.
  • To derive asymptotic expansions of critical points in powers of $k^{-1/2}$, where $k$ is the number of hidden neurons.
  • To quantify differences in loss values at spurious minima and show they are not uniformly small.
  • To establish that some spurious minima are fundamentally worse than others due to persistent non-zero loss.

Proposed method

  • Exploiting the inherent symmetry of ReLU networks with k neurons to identify families of critical points via group action geometry.
  • Applying Artin's implicit function theorem to construct local parameterizations of critical points as power series in $k^{-1/2}$.
  • Using bifurcation theory to analyze how critical points emerge and evolve with increasing $k$.
  • Deriving asymptotic expressions for critical points and their associated loss values in the large-$k$ regime.
  • Analyzing the loss function at critical points to compare decay rates across different types of spurious minima.
  • Leveraging symmetry to reduce the dimensionality of the optimization landscape and isolate key critical configurations.

Experimental results

Research questions

  • RQ1How do the critical points of a two-layer ReLU network with k neurons depend on the underlying symmetry of the model?
  • RQ2What is the asymptotic behavior of critical points as $k \to \infty$, and how can they be parameterized?
  • RQ3Do all spurious minima lead to similar loss values, or do they differ significantly in optimization quality?
  • RQ4Can the loss at spurious minima be characterized by distinct decay rates, and if so, what determines these rates?
  • RQ5What role does the geometry of group actions play in organizing the critical point structure of shallow ReLU networks?

Key findings

  • Critical points in the two-layer ReLU network are classified into distinct families based on symmetry, with each family admitting a power series expansion in $k^{-1/2}$.
  • The loss at certain spurious minima decays to zero as $k^{-1}$, indicating favorable optimization properties in the large-$k$ limit.
  • In contrast, other spurious minima exhibit a loss that converges to a strictly positive constant, indicating persistent optimization error.
  • These results demonstrate that spurious minima are not uniformly problematic—some are significantly worse than others in terms of final loss.
  • The use of symmetry and Artin's implicit function theorem enables rigorous asymptotic analysis of critical points in non-convex, high-dimensional optimization landscapes.
  • The geometry of group actions provides a structural framework for understanding the organization and diversity of critical points in shallow neural networks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.