Skip to main content
QUICK REVIEW

[论文解读] Optimal subsampling for quantile regression in big data

HaiYing Wang, Yanyuan Ma|arXiv (Cornell University)|Jan 28, 2020
Statistical Methods and Inference参考文献 20被引用 5
一句话总结

本文提出了大数据中分位数回归的最优子采样方法,通过两种最优概率版本最小化渐近方差——一种与响应密度无关且易于实现。迭代子采样过程确保了计算可扩展性,并在无需密度估计的情况下实现标准误估计,达到渐近最优性,并在大规模数据集上实现稳健推断。

ABSTRACT

We investigate optimal subsampling for quantile regression. We derive the asymptotic distribution of a general subsampling estimator and then derive two versions of optimal subsampling probabilities. One version minimizes the trace of the asymptotic variance-covariance matrix for a linearly transformed parameter estimator and the other minimizes that of the original parameter estimator. The former does not depend on the densities of the responses given covariates and is easy to implement. Algorithms based on optimal subsampling probabilities are proposed and asymptotic distributions and asymptotic optimality of the resulting estimators are established. Furthermore, we propose an iterative subsampling procedure based on the optimal subsampling probabilities in the linearly transformed parameter estimation which has great scalability to utilize available computational resources. In addition, this procedure yields standard errors for parameter estimators without estimating the densities of the responses given the covariates. We provide numerical examples based on both simulated and real data to illustrate the proposed method.

研究动机与目标

  • 解决在传统优化方法过于缓慢的大规模数据集中分位数回归的计算负担问题。
  • 克服由于给定协变量下响应密度未知而导致的分位数回归渐近方差-协方差矩阵估计的挑战。
  • 推导最小化分位数回归估计量渐近均方误差的最优子采样概率。
  • 提出一种迭代子采样算法,提升可扩展性并实现在无需密度估计情况下的标准误估计。
  • 在所提出的子采样框架下,建立结果估计量的渐近分布及其最优性理论。

提出的方法

  • 推导在数据和子采样双重随机性来源下,一般子采样估计量的渐近分布。
  • 提出两种最优子采样概率版本:一种是最小化原始参数估计量渐近协方差矩阵的迹,另一种是针对线性变换参数估计量的最优概率。
  • 后一种版本无需依赖给定协变量下响应的条件密度,从而实现更简便的实现。
  • 基于变换参数的最优概率设计一种迭代子采样过程,提升可扩展性与计算效率。
  • 利用该迭代过程在不估计给定协变量下响应密度的情况下估计标准误,简化推断过程。
  • 在所提出的子采样方案下,建立结果估计量的渐近正态性与渐近最优性。

实验结果

研究问题

  • RQ1如何为大数据中的分位数回归推导最优子采样概率,以最小化渐近均方误差?
  • RQ2子采样能否在保持统计效率的同时实现可扩展性与计算高效性,并支持推断?
  • RQ3如何在不依赖给定协变量下响应密度估计的情况下,估计分位数回归的标准误?
  • RQ4在最优子采样概率下,子采样估计量的渐近分布是什么?
  • RQ5与分治法等替代方法相比,迭代子采样过程在计算时间与估计精度方面表现如何?

主要发现

  • 所提方法通过在最优子采样概率下最小化渐近协方差矩阵的迹,实现了渐近最优性。
  • 针对线性变换参数估计量的最优概率版本不依赖于给定协变量下响应的未知密度,从而实现简便实现。
  • 数值研究显示,即使在子样本量较小(n₀ = 1000)且全量数据较大(N = 10⁶)时,95%置信区间的覆盖概率仍接近名义水平(如~95%),表明推断可靠。
  • 当B = 100或500时,方法在τ = 0.5和τ = 0.75下保持良好覆盖,但当n₀ = 100且B ≥ 100时,覆盖率略有下降(如β₁在τ = 0.75时约为~83%),提示在子样本量较小时需谨慎。
  • 尽管存在计算最优概率的开销,迭代子采样过程仍显著快于分治法,因其依赖于小样本而非全量数据块。
  • 经验均方误差随计算时间增加而减小,且所提方法在均方误差与时间的权衡空间中占据与分治法不同的区域,表明其性能特征因计算资源与精度优先级的不同而异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。