[论文解读] Guarding from Spurious Discoveries in High Dimension
本文提出了一种广义的最大虚假相关性度量,通过LAMM算法计算,用于量化在零假设模型下,响应变量能被随机协变量子集拟合的程度。该文推导了广义线性模型和L1-回归下该度量的渐近分布,从而通过乘子自展法实现一致的基准比较,以防范高维数据中的虚假发现。
Many data-mining and statistical machine learning algorithms have been developed to select a subset of covariates to associate with a response variable. Spurious discoveries can easily arise in high-dimensional data analysis due to enormous possibilities of such selections. How can we know statistically our discoveries better than those by chance? In this paper, we define a measure of goodness of spurious fit, which shows how good a response variable can be fitted by an optimally selected subset of covariates under the null model, and propose a simple and effective LAMM algorithm to compute it. It coincides with the maximum spurious correlation for linear models and can be regarded as a generalized maximum spurious correlation. We derive the asymptotic distribution of such goodness of spurious fit for generalized linear models and $L_1$-regression. Such an asymptotic distribution depends on the sample size, ambient dimension, the number of variables used in the fit, and the covariance information. It can be consistently estimated by multiplier bootstrapping and used as a benchmark to guard against spurious discoveries. It can also be applied to model selection, which considers only candidate models with goodness of fits better than those by spurious fits. The theory and method are convincingly illustrated by simulated examples and an application to the binary outcomes from German Neuroblastoma Trials.
研究动机与目标
- 为解决高维数据分析中普遍存在的虚假发现问题,即大量协变量选择可能错误地表现为预测性。
- 定义一种统计上严谨的‘虚假拟合优度’度量,用于量化在零假设下,随机协变量子集所能达到的最佳拟合程度。
- 开发一种计算高效的算法(LAMM),用于计算该虚假拟合度量,将最大虚假相关性的概念从线性模型推广至更广范围。
- 推导广义线性模型和L1-回归下虚假拟合优度的渐近分布,其依赖于样本量、维度和设计协方差结构。
- 通过拒绝那些拟合度不显著优于基准虚假拟合的模型候选,实现模型选择,从而防范虚假发现。
提出的方法
- 提出广义的最大虚假相关性作为虚假拟合优度的度量,将适用范围从线性模型扩展至广义线性模型和L1-正则化回归。
- 引入LAMM算法(拉格朗日增强最大化法),以高效计算零假设模型下的最大虚假相关性。
- 推导虚假拟合优度的渐近分布,其依赖于样本量、环境维度、拟合中使用的协变量数量以及设计协方差结构。
- 采用乘子自展法一致估计渐近分布,从而实现统计显著性阈值的经验校准。
- 将该基准应用于模型选择,通过剔除拟合度不显著优于虚假拟合阈值的模型,实现筛选。
- 通过模拟实验和对德国神经母细胞瘤试验二元结果的真实应用验证该方法。
实验结果
研究问题
- RQ1当预测变量与响应变量之间不存在真实关系时,如何统计量化随机协变量子集所能达到的最佳拟合程度?
- RQ2在高维设置下,广义线性模型和L1-回归中最大虚假拟合的渐近分布是什么?
- RQ3能否使用乘子自展法一致估计虚假拟合的渐近分布,以实现实际推断?
- RQ4如何利用虚假拟合的基准来改进模型选择,通过剔除虚假发现?
- RQ5该方法在真实高维数据中检测虚假关联的实证表现如何?
主要发现
- 所提出的虚假拟合优度度量将最大虚假相关性从线性模型推广至更广范围,为高维模型评估提供了统一框架。
- 虚假拟合优度的渐近分布依赖于样本量、环境维度、拟合中使用的协变量数量以及设计协方差矩阵。
- 可通过乘子自展法一致估计渐近分布,从而实现方法的实际应用与推断。
- 基于虚假拟合分布推导出的基准能有效防范虚假发现,通过设定模型选择的统计阈值实现。
- 模拟实验和对德国神经母细胞瘤试验数据的应用表明,该方法在高维设置下能有效区分真实信号与虚假相关性。
- LAMM算法能高效计算最大虚假相关性,使该方法具备可扩展性,适用于真实世界数据分析。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。