[Paper Review] Optimal Subsampling Algorithms for Big Data Generalized Linear Models
This paper extends the Optimal Subsampling Method under the A-optimality Criterion (OSMAC) to generalized linear models (GLMs) with canonical links. It derives optimal subsampling probabilities under A- and L-optimality, establishes asymptotic normality and consistency, and proposes an adaptive two-step algorithm that achieves asymptotic optimality and finite-sample efficiency, validated through simulations and real data.
To fast approximate the maximum likelihood estimator with massive data, Wang et al. (JASA, 2017) proposed an Optimal Subsampling Method under the A-optimality Criterion (OSMAC) for in logistic regression. This paper extends the scope of the OSMAC framework to include generalized linear models with canonical link functions. The consistency and asymptotic normality of the estimator from a general subsampling algorithm are established, and optimal subsampling probabilities under the A- and L-optimality criteria are derived. Furthermore, using Frobenius norm matrix concentration inequality, finite sample properties of the subsample estimator based on optimal subsampling probabilities are derived. Since the optimal subsampling probabilities depend on the full data estimate, an adaptive two-step algorithm is developed. Asymptotic normality and optimality of the estimator from this adaptive algorithm are established. The proposed methods are illustrated and evaluated through numerical experiments on simulated and real datasets.
Motivation & Objective
- Address the computational burden of fitting generalized linear models (GLMs) on massive datasets by developing efficient subsampling methods.
- Extend the OSMAC framework—previously limited to logistic regression—to general GLMs with canonical link functions.
- Establish theoretical properties such as consistency and asymptotic normality for subsample estimators under general subsampling algorithms.
- Derive optimal subsampling probabilities under both A- and L-optimality criteria for GLMs.
- Develop and analyze an adaptive two-step algorithm that achieves asymptotic optimality and finite-sample efficiency.
Proposed method
- Derive optimal subsampling probabilities under A- and L-optimality criteria for GLMs with canonical link functions using asymptotic variance-covariance approximations.
- Use Frobenius norm matrix concentration inequalities to establish finite-sample properties of the subsample estimator under optimal probabilities.
- Propose an adaptive two-step algorithm: first estimate full-data parameters, then compute optimal subsampling probabilities based on these estimates.
- Ensure the resulting estimator is consistent and asymptotically normal under the proposed adaptive scheme.
- Apply the subsampling algorithm to both simulated and real datasets to evaluate performance empirically.
- Leverage the asymptotic distributional properties of the MLE in GLMs to justify the optimality criteria and estimator behavior.
Experimental results
Research questions
- RQ1Can the OSMAC framework be generalized to handle generalized linear models beyond logistic regression?
- RQ2What are the optimal subsampling probabilities under A- and L-optimality criteria for GLMs with canonical links?
- RQ3Do subsample estimators based on optimal probabilities achieve consistency and asymptotic normality in the GLM setting?
- RQ4Can an adaptive two-step algorithm be designed to achieve asymptotic optimality when full-data estimates are unknown?
- RQ5How do the finite-sample properties of the subsample estimator compare to the full-data MLE under optimal subsampling?
Key findings
- The proposed optimal subsampling probabilities under A- and L-optimality criteria are derived for GLMs with canonical link functions.
- The subsample estimator based on optimal probabilities is consistent and asymptotically normal under regularity conditions.
- Finite-sample properties of the estimator are established using Frobenius norm matrix concentration inequalities.
- The adaptive two-step algorithm achieves asymptotic normality and optimality, even when full-data estimates are used to construct probabilities.
- Numerical experiments on simulated and real datasets demonstrate the method's efficiency and accuracy relative to full-data MLE.
- The method significantly reduces computational cost while maintaining high statistical efficiency in large-scale GLM settings.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.