Skip to main content
QUICK REVIEW

[论文解读] Method of Divide-and-Combine in Regularised Generalised Linear Models for Big Data

Lu Tang, Ling Zhou|arXiv (Cornell University)|Nov 18, 2016
Advanced Statistical Methods and Models参考文献 25被引用 17
一句话总结

本文提出了一种用于大规模数据中正则化广义线性模型的新型分治-合并方法,利用基于偏差校正估计量的置信分布来合并套索型回归估计。该方法实现了费雪效率——其估计精度与全数据最大似然估计相当——同时实现了可扩展且稳定的变量选择。

ABSTRACT

When a data set is too big to be analysed entirely once by a single computer, the strategy of divide-and-combine has been the method of choice to overcome the computational hurdle due to its scalability. Although random data partition has been widely adopted, there is lack of clear theoretical justification and practical guidelines to combine results obtained from separate analysis of individual sub-datasets, especially when a regularisation method such as lasso is utilised for variable selection to improve numerical stability. In this paper we develop a new strategy to combine separate lasso-type estimates of regression parameters by the means of the confidence distributions based on bias-corrected estimators. We first establish the approach to the construction of the confidence distribution and then show that the resulting combined estimator enjoys the Fisher's efficiency in the sense of the estimation efficiency achieved by the maximum likelihood estimator from the analysis of full data. Furthermore, using the combined regularised estimator we propose an inference procedure. Extensive simulation studies are presented to evaluate the performance of the proposed methodology with comparisons to the classical meta estimation method and a voting-based variable selection method.

研究动机与目标

  • 解决在大规模数据中使用正则化(如套索)时,对子数据集结果合并缺乏理论依据和实际指导的问题。
  • 开发一种可扩展、高效的合并方法,无需处理完整数据集即可合并分割数据中的独立套索估计。
  • 确保合并后的估计量达到与全数据集分析中最大似然估计相当的估计效率。
  • 基于合并后的正则化估计量提出有效的推断程序,用于统计检验和区间估计。

提出的方法

  • 利用各子数据集中的偏差校正估计量构建置信分布,以校正数据分割引入的估计偏差。
  • 通过加权平均方案合并子数据集的置信分布,以保持估计效率。
  • 利用合并后的置信分布推导出最终的正则化估计量,使其估计效率与全数据集最大似然估计量相当。
  • 将合并后的估计量应用于推断,包括假设检验和置信区间的构建,均在正则化模型框架下进行。
  • 通过避免全数据计算,在保持统计效率的同时确保数值稳定性和可扩展性。

实验结果

研究问题

  • RQ1在广义线性模型中,针对套索型估计的分治-合并策略是否能在不处理整个数据集的情况下实现费雪效率?
  • RQ2如何构建并合并置信分布,以从子数据集中获得可靠且高效的估计量?
  • RQ3所提出的方法在估计精度和变量选择一致性方面是否优于经典元估计和基于投票的变量选择方法?
  • RQ4合并估计量在数据分割和正则化条件下的效率具有何种理论依据?

主要发现

  • 所提出的合并估计量实现了费雪效率,即其估计精度与从全数据集中推导出的最大似然估计量相当。
  • 该方法为从子数据集合并套索估计提供了一种理论基础坚实的方法,克服了现有分治-合并策略中缺乏理论依据的问题。
  • 模拟研究结果表明,所提出方法在变量选择准确性和估计精度方面优于经典元估计和基于投票的变量选择方法。
  • 基于合并估计量的推断程序在置信区间构造中实现了有效的覆盖率,且在假设检验中保持了适当的 I 类错误率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。