Skip to main content
QUICK REVIEW

[论文解读] Two-sample testing in non-sparse high-dimensional linear models

Yinchu Zhu, Jelena Bradić|arXiv (Cornell University)|Oct 14, 2016
Statistical Methods and Inference参考文献 5被引用 6
一句话总结

该论文提出TIERS,一种针对高维线性模型的新型两样本检验框架,无需假设回归系数的稀疏性。通过卷积两个样本,将斜率相等的原假设转化为矩条件,并利用自标准化方法,在弱矩条件和尾部条件下实现稳健的 I 类错误控制,即使当 $ p \to \infty $ 且 $ n \to \infty $,且 $ p/n \to \infty $ 时亦然,同时引入了自动自适应Dantzig选择器(ADDS)以及一种高维插补临界值近似方法。

ABSTRACT

In analyzing high-dimensional models, sparsity of the model parameter is a common but often undesirable assumption. In this paper, we study the following two-sample testing problem: given two samples generated by two high-dimensional linear models, we aim to test whether the regression coefficients of the two linear models are identical. We propose a framework named TIERS (short for TestIng Equality of Regression Slopes), which solves the two-sample testing problem without making any assumptions on the sparsity of the regression parameters. TIERS builds a new model by convolving the two samples in such a way that the original hypothesis translates into a new moment condition. A self-normalization construction is then developed to form a moment test. We provide rigorous theory for the developed framework. Under very weak conditions of the feature covariance, we show that the accuracy of the proposed test in controlling Type I errors is robust both to the lack of sparsity in the features and to the heavy tails in the error distribution, even when the sample size is much smaller than the feature dimension. Moreover, we discuss minimax optimality and efficiency properties of the proposed test. Simulation analysis demonstrates excellent finite-sample performance of our test. In deriving the test, we also develop tools that are of independent interest. The test is built upon a novel estimator, called Auto-aDaptive Dantzig Selector (ADDS), which not only automatically chooses an appropriate scale of the error term but also incorporates prior information. To effectively approximate the critical value of the test statistic, we develop a novel high-dimensional plug-in approach that complements the recent advances in Gaussian approximation theory.

研究动机与目标

  • 为解决在回归系数非稀疏时,高维两样本检验缺乏系统性推断方法的问题。
  • 开发一种在特征协方差和误差分布假设较弱(即使存在重尾)时仍能保持准确 I 类错误控制的检验方法。
  • 构建一种在 $ p \to \infty $、$ n \to \infty $ 且 $ p/n \to \infty $ 情况下仍稳健的方法,且无需真实回归参数具有稀疏性的要求。
  • 提出一种新型估计器——自动自适应Dantzig选择器(ADDS),可自适应选择误差方差尺度并整合先验信息。
  • 开发一种高维插补方法,用于近似检验统计量的临界值,补充了近期的高斯近似理论。

提出的方法

  • TIERS通过构造两个样本的卷积,将原始两样本假设 $ H_0: \beta_A = \beta_B $ 转化为新的矩条件。
  • 对变换后的数据应用自标准化程序,形成在原假设下渐近分布自由的枢轴检验统计量。
  • 引入自动自适应Dantzig选择器(ADDS)作为新型估计器,可自适应选择误差方差尺度并整合先验知识。
  • 使用高维插补方法近似检验统计量的临界值,提升有限样本性能,并避免使用自助法或重抽样。
  • 理论分析依赖于反浓度不等式和次高斯尾部界,以控制高维随机向量最大值的偏离。
  • 利用关于依赖随机向量最大值收敛性及条件高斯性下的反浓度性质的引理,建立检验的稳健性。

实验结果

研究问题

  • RQ1我们能否为高维线性模型开发一种不依赖于回归系数稀疏性假设的两样本检验程序?
  • RQ2当特征数 $ p $ 远大于样本量 $ n $ 时,即使在非稀疏模型下,如何确保 I 类错误控制的准确性?
  • RQ3重尾误差和非稀疏设计矩阵对高维两样本检验有效性有何影响?
  • RQ4我们能否构建一种检验统计量,其临界值可通过插补方法在高维设置下近似,而无需重抽样?
  • RQ5所提出的检验在非稀疏高维模型下是否达到极小极大最优性或效率?

主要发现

  • TIERS在特征协方差结构条件极弱时仍能准确控制 I 类错误,即使真实回归系数为稠密。
  • 该检验对重尾误差分布具有鲁棒性,无需假设次高斯或轻尾分布。
  • 自动自适应Dantzig选择器(ADDS)可自动选择误差方差尺度并整合先验信息,提升估计稳定性。
  • 高维插补方法在临界值近似中表现出优异的有限样本性能,且避免了计算密集的重抽样过程。
  • 理论分析表明,该检验对回归参数的非稀疏性以及非稀疏设计矩阵均具有鲁棒性,即使在 $ p/n \to \infty $ 时亦然。
  • 通过模拟研究证明,该方法在适当的正则化条件下实现了极小极大最优性和效率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。