Skip to main content
QUICK REVIEW

[论文解读] Guarding against Spurious Discoveries in High Dimensions

Jianqing Fan, Wen‐Xin Zhou|PubMed|Dec 5, 2015
Statistical Methods and Inference参考文献 1被引用 6
一句话总结

本文提出一种广义的最大虚假相关性度量,称为虚假拟合优度(GOSF),用于评估高维变量选择结果是否显著优于随机机会。提出LAMM算法以计算GOSF,并推导其在广义线性模型和$L_1$回归下的渐近分布,从而通过乘子自展法实现一致的基准比较,用于模型选择与虚假发现控制。

ABSTRACT

Many data-mining and statistical machine learning algorithms have been developed to select a subset of covariates to associate with a response variable. Spurious discoveries can easily arise in high-dimensional data analysis due to enormous possibilities of such selections. How can we know statistically our discoveries better than those by chance? In this paper, we define a measure of goodness of spurious fit, which shows how good a response variable can be fitted by an optimally selected subset of covariates under the null model, and propose a simple and effective LAMM algorithm to compute it. It coincides with the maximum spurious correlation for linear models and can be regarded as a generalized maximum spurious correlation. We derive the asymptotic distribution of such goodness of spurious fit for generalized linear models and <i>L</i><sub>1</sub>-regression. Such an asymptotic distribution depends on the sample size, ambient dimension, the number of variables used in the fit, and the covariance information. It can be consistently estimated by multiplier bootstrapping and used as a benchmark to guard against spurious discoveries. It can also be applied to model selection, which considers only candidate models with goodness of fits better than those by spurious fits. The theory and method are convincingly illustrated by simulated examples and an application to the binary outcomes from German Neuroblastoma Trials.

研究动机与目标

  • 为解决高维变量选择中虚假发现的关键问题,即所选协变量可能仅因偶然性而表现出预测能力。
  • 定义一个严格的统计基准——虚假拟合优度(GOSF),以量化在无真实关联的原假设下,任意协变量子集所能达到的最佳拟合程度。
  • 将最大虚假相关性的概念从线性模型扩展至广义线性模型和$L_1$回归。
  • 推导GOSF的渐近分布,其依赖于样本量、可观测维度、模型大小及设计协方差结构。
  • 提供一种实用且一致的方法,通过乘子自展法估计GOSF分布,从而实现排除虚假拟合的模型选择。

提出的方法

  • 提出虚假拟合优度(GOSF)作为在原模型下,任意$s$-子集协变量所能达到的最大似然比统计量的度量。
  • 提出LAMM算法,通过在原假设下对所有$s$-子集协变量进行优化,高效计算GOSF。
  • 利用非渐近的、条件化的威尔克斯定理版本,推导广义线性模型和$L_1$回归下GOSF的渐近分布。
  • 证明在正则性条件下,所有$s$-子集上,对数似然的超额平方根与归一化得分范数在一致意义上接近。
  • 使用归一化得分$\widehat{\boldsymbol{\xi}}_S$作为超额对数似然的代理,其$\ell_2$-范数近似GOSF。
  • 应用乘子自展法一致估计GOSF的渐近分布,从而实现对实际模型拟合的实证基准比较。

实验结果

研究问题

  • RQ1当所有协变量与响应变量真正无关时,如何正式度量响应变量通过任意协变量子集所能达到的最佳拟合?
  • RQ2在高维设定下,广义线性模型和$L_1$回归中最大虚假拟合的渐近分布为何?
  • RQ3在实践中,虚假拟合优度能否被一致估计,以作为模型选择的基准?
  • RQ4GOSF分布如何依赖于样本量、可观测维度和模型大小等关键设计参数?
  • RQ5所提出的方法在多大程度上可防范高维数据分析中的虚假发现?

主要发现

  • 虚假拟合优度(GOSF)的渐近分布依赖于样本量$n$、可观测维度$p$、模型中变量数$s$以及设计矩阵的协方差结构。
  • 可通过乘子自展法一致估计GOSF分布,从而实现在高维模型选择中对统计显著性的实证校准。
  • LAMM算法通过在所有$s$-子集协变量上进行优化,高效计算GOSF,并在正则性条件下具备理论保证。
  • 归一化得分$\widehat{\boldsymbol{\xi}}_S$在所有$s$-子集上以高概率接近对数似然超额的平方根。
  • 理论逼近误差界$\|\widehat{\boldsymbol{\xi}}_S\|_2$与$\|\boldsymbol{\xi}_S\|_2$的阶为$O(s \log(pn) / n^{1/2})$,确保在高维情形下的稳健性。
  • 该方法在模拟数据上得到验证,并应用于德国神经母细胞瘤试验的二值结果,展示了其在真实高维场景中的实际效用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。