[Paper Review] Rates of Convergence for Laplacian Semi-Supervised Learning with Low Labeling Rates
This paper analyzes the convergence rates of Laplacian semi-supervised learning under low labeling rates using graph-based regularization. It establishes that degeneracy (spikes at labeled points) occurs when labeling rate β ≪ ε², while consistency with the continuum Laplace equation holds when β ≫ ε², with error bounds of O(εβ⁻¹ᐟ²) up to logarithmic factors, using random walk and Γ-convergence tools for proofs.
We study graph-based Laplacian semi-supervised learning at low labeling rates. Laplacian learning uses harmonic extension on a graph to propagate labels. At very low label rates, Laplacian learning becomes degenerate and the solution is roughly constant with spikes at each labeled data point. Previous work has shown that this degeneracy occurs when the number of labeled data points is finite while the number of unlabeled data points tends to infinity. In this work we allow the number of labeled data points to grow to infinity with the number of labels. Our results show that for a random geometric graph with length scale $\varepsilon>0$ and labeling rate $β>0$, if $β\ll\varepsilon^2$ then the solution becomes degenerate and spikes form, and if $β\gg \varepsilon^2$ then Laplacian learning is well-posed and consistent with a continuum Laplace equation. Furthermore, in the well-posed setting we prove quantitative error estimates of $O(\varepsilonβ^{-1/2})$ for the difference between the solutions of the discrete problem and continuum PDE, up to logarithmic factors. We also study $p$-Laplacian regularization and show the same degeneracy result when $β\ll \varepsilon^p$. The proofs of our well-posedness results use the random walk interpretation of Laplacian learning and PDE arguments, while the proofs of the ill-posedness results use $Γ$-convergence tools from the calculus of variations. We also present numerical results on synthetic and real data to illustrate our results.
Motivation & Objective
- To understand the asymptotic behavior of Laplacian semi-supervised learning when the number of labeled points grows with the total data size.
- To identify the critical threshold between degenerate and well-posed regimes in graph-based learning under low labeling rates.
- To establish quantitative error estimates for the convergence of discrete graph solutions to the continuum PDE solution.
- To extend the analysis to p-Laplacian regularization and derive analogous degeneracy conditions.
- To validate theoretical findings with numerical experiments on synthetic and real data.
Proposed method
- Uses a random geometric graph with edge weights defined by a kernel ηε(|xi−xj|) to model data structure.
- Analyzes Laplacian learning via the graph Laplace equation and its equivalence to harmonic extension and random walk expectations.
- Applies the random walk interpretation to derive probabilistic intuition for label propagation and convergence.
- Employs Γ-convergence techniques from the calculus of variations to prove well-posedness and consistency in the limit.
- Derives error bounds using PDE arguments and measure-theoretic convergence tools, including Lp convergence and transport maps.
- Extends results to p-Laplacian regularization by analyzing nonlocal Dirichlet energies and their Γ-limits.
Experimental results
Research questions
- RQ1Under what conditions does Laplacian semi-supervised learning become degenerate at low labeling rates?
- RQ2What is the critical scaling between labeling rate β and graph length scale ε that separates degenerate from consistent behavior?
- RQ3What is the rate of convergence of the discrete graph solution to the continuum PDE solution in the well-posed regime?
- RQ4How does the analysis extend to p-Laplacian regularization, and what is the corresponding degeneracy threshold?
- RQ5Can the theoretical convergence results be validated through numerical experiments on synthetic and real data?
Key findings
- Degeneracy occurs when β ≪ ε², leading to solutions that are nearly constant with sharp spikes at labeled points.
- Well-posedness and consistency with the continuum Laplace equation hold when β ≫ ε².
- The error between the discrete solution and the continuum PDE solution is bounded by O(εβ⁻¹ᐟ²), up to logarithmic factors.
- For p-Laplacian regularization, degeneracy occurs when β ≪ εᵖ, generalizing the quadratic case.
- Theoretical convergence is supported by numerical experiments on synthetic and real datasets, confirming the predicted scaling behavior.
- Γ-convergence and random walk interpretations provide rigorous justification for the asymptotic behavior and error estimates.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.