[Paper Review] Convex and Non-convex Approaches for Statistical Inference with Class-Conditional Noisy Labels
This paper proposes convex and non-convex estimation approaches for logistic regression under class-conditional label noise, modeling the problem as a generalized linear model with latent true labels. It establishes asymptotic normality and minimax optimal $β$-consistency in high dimensions, with the non-convex maximum likelihood estimator achieving lower asymptotic variance and global convergence under mild conditions.
We study the problem of estimation and testing in logistic regression with class-conditional noise in the observed labels, which has an important implication in the Positive-Unlabeled (PU) learning setting. With the key observation that the label noise problem belongs to a special sub-class of generalized linear models (GLM), we discuss convex and non-convex approaches that address this problem. A non-convex approach based on the maximum likelihood estimation produces an estimator with several optimal properties, but a convex approach has an obvious advantage in optimization. We demonstrate that in the low-dimensional setting, both estimators are consistent and asymptotically normal, where the asymptotic variance of the non-convex estimator is smaller than the convex counterpart. We also quantify the efficiency gap which provides insight into when the two methods are comparable. In the high-dimensional setting, we show that both estimation procedures achieve $\\ell_2$-consistency at the minimax optimal $\\sqrt{s\\log p/n}$ rates under mild conditions. Finally, we propose an inference procedure using a de-biasing approach. We validate our theoretical findings through simulations and a real-data example.
Motivation & Objective
- Address statistical inference in logistic regression when observed labels are corrupted by class-conditional noise, a key challenge in Positive-Unlabeled (PU) learning.
- Bridge the gap between theoretical optimality and computational feasibility by comparing convex and non-convex estimation strategies.
- Establish consistency and asymptotic normality in low dimensions and $β$-consistency at minimax optimal rates in high dimensions.
- Develop a de-biasing procedure to enable valid inference (e.g., confidence intervals, p-values) under label noise.
- Provide theoretical guarantees on global convergence for the non-convex likelihood-based estimator, overcoming local optima issues common in prior work.
Proposed method
- Model label noise as a class-conditional contamination process where true labels are stochastically corrupted with asymmetric probabilities independent of features.
- Formulate the problem as a generalized linear model (GLM) with latent true labels, enabling likelihood-based inference.
- Propose a non-convex maximum likelihood estimator (MLE) that achieves global convergence with high probability under mild regularity conditions.
- Develop a convex relaxation using a surrogate loss function to ensure computational tractability and efficient optimization.
- Apply a de-biasing technique to the regularized estimators to construct asymptotically normal estimators for valid inference.
- Use node-wise lasso regression to construct an approximate inverse of the Fisher information matrix for de-biasing, leveraging high-dimensional concentration inequalities.
Experimental results
Research questions
- RQ1Can non-convex maximum likelihood estimation achieve global convergence and optimal statistical efficiency under class-conditional label noise?
- RQ2How do the asymptotic variances of convex and non-convex estimators compare in low-dimensional settings?
- RQ3What are the high-dimensional estimation rates for both convex and non-convex approaches under $β$-regularization?
- RQ4Can a de-biasing procedure yield asymptotically normal estimators that support valid inference (e.g., confidence intervals) under label noise?
- RQ5What is the efficiency gap between convex and non-convex estimators, and when are they comparable in practice?
Key findings
- In low dimensions, both convex and non-convex estimators are consistent and asymptotically normal, with the non-convex estimator achieving a smaller asymptotic variance.
- The efficiency gap between the two estimators is quantified, showing that the non-convex approach is strictly more efficient under the same conditions.
- In high dimensions, both estimators achieve $β$-consistency at the minimax optimal rate of $\sqrt{s\log p/n}$ under mild regularity conditions.
- The de-biasing procedure yields asymptotically normal estimators, enabling valid inference such as confidence intervals and hypothesis tests.
- The non-convex MLE achieves global convergence with high probability, resolving a key limitation in prior work where global optima were not guaranteed.
- The approximate inverse of the Fisher information matrix constructed via node-wise lasso is consistent in $\ell_1$ and $\ell_2$ norms, supporting the de-biasing framework.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.