[Paper Review] Pure and Spurious Critical Points: a Geometric Study of Linear Networks
This paper introduces a geometric framework distinguishing pure critical points—arising from the functional space of linear networks—from spurious critical points—caused by parameterization. It proves that for filling architectures (capable of expressing all linear maps), any smooth convex loss has no bad local minima, while for non-filling architectures (with rank-constrained functional spaces), only the quadratic loss guarantees no bad minima due to special geometric properties of determinantal varieties.
The critical locus of the loss function of a neural network is determined by the geometry of the functional space and by the parameterization of this space by the network's weights. We introduce a natural distinction between pure critical points, which only depend on the functional space, and spurious critical points, which arise from the parameterization. We apply this perspective to revisit and extend the literature on the loss function of linear neural networks. For this type of network, the functional space is either the set of all linear maps from input to output space, or a determinantal variety, i.e., a set of linear maps with bounded rank. We use geometric properties of determinantal varieties to derive new results on the landscape of linear networks with different loss functions and different parameterizations. Our analysis clearly illustrates that the absence of "bad" local minima in the loss landscape of linear networks is due to two distinct phenomena that apply in different settings: it is true for arbitrary smooth convex losses in the case of architectures that can express all linear maps ("filling architectures") but it holds only for the quadratic loss when the functional space is a determinantal variety ("non-filling architectures"). Without any assumption on the architecture, smooth convex losses may lead to landscapes with many bad minima.
Motivation & Objective
- To resolve the longstanding puzzle of why linear networks often avoid bad local minima despite non-convexity.
- To formally distinguish between critical points arising from the functional space (pure) and those from parameterization (spurious).
- To analyze the loss landscape of linear networks through algebraic geometry, particularly determinantal varieties.
- To clarify under which conditions smooth convex losses yield no non-global minima in linear networks.
- To unify prior results on linear network optimization by identifying two distinct geometric mechanisms for absence of bad minima.
Proposed method
- Introduces a decomposition of the loss function as a composition: parameter space → functional space → R, with the functional space being a set of linear maps.
- Defines pure critical points as those determined solely by the geometry of the functional space, and spurious ones as artifacts of the parameterization map.
- Analyzes the differential of matrix multiplication to characterize critical points in linear networks.
- Uses algebraic geometry tools to study determinantal varieties (rank-constrained linear maps), particularly their singularities and curvature.
- Applies Schur complements and eigenvalue analysis to compute the Hessian’s characteristic polynomial and count negative eigenvalues.
- Derives conditions under which the Hessian has no negative eigenvalues (i.e., no bad local minima) via explicit computation of the characteristic polynomial.
Experimental results
Research questions
- RQ1What causes the absence of non-global local minima in linear networks, and is this phenomenon universal across all loss functions?
- RQ2How do the geometric properties of the functional space (e.g., determinantal varieties) influence the loss landscape?
- RQ3In what settings do spurious critical points dominate, and when are they absent?
- RQ4Why does the quadratic loss avoid bad minima in non-filling architectures, while other convex losses do not?
- RQ5Can the distinction between pure and spurious critical points explain prior results on linear network optimization?
Key findings
- For filling architectures (where the network can express all linear maps), any smooth convex loss function has no bad local minima.
- For non-filling architectures (where the functional space is a determinantal variety), only the quadratic loss guarantees no bad local minima.
- The absence of bad local minima in the quadratic loss case is due to special geometric properties of determinantal varieties, not general convexity.
- For arbitrary smooth convex losses on non-filling architectures, the loss landscape can contain many non-global local minima.
- The number of negative eigenvalues of the Hessian (indicating bad local minima) is determined by the relative singular values of the input and output layers.
- The characteristic polynomial of the Hessian is explicitly computed, and its negative roots (corresponding to bad local minima) are counted via algebraic analysis of the singular values.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.