[论文解读] A Scale-Free Approach for False Discovery Rate Control in Generalized Linear Models
本论文提出了一种用于广义线性模型(GLMs)的无尺度错误发现率(FDR)控制框架,通过数据分割或高斯镜像法从两个渐近独立的系数估计中推导出镜像统计量。该方法利用镜像统计量在原假设下对称的抽样分布,实现FDR控制,且在中等维和高维设置下均优于Benjamini-Hochberg和knockoff过滤器,而无需依赖渐近正态性。
The Generalized Linear Model (GLM) has been widely used in practice to model counts or other types of non-Gaussian data. This article introduces a framework for feature selection in the GLM that can achieve robust False Discovery Rate (FDR) control. The main idea is to construct a <i>mirror statistic</i> based on data perturbation to measure the importance of each feature. FDR control is achieved by taking advantage of the mirror statistic’s property that its sampling distribution is (asymptotically) symmetric about zero for any null feature. In the moderate-dimensional setting, that is, p/n→κ∈(0,1), we construct the mirror statistic based on the maximum likelihood estimation. In the high-dimensional setting, that is, p≫n, we use the debiased Lasso to build the mirror statistic. The proposed methodology is scale-free as it only hinges on the symmetry of the mirror statistic, thus, can be more robust in finite-sample cases compared to existing methods. Both simulation results and a real data application show that the proposed methods are capable of controlling the FDR and are often more powerful than existing methods including the Benjamini-Hochberg procedure and the knockoff filter. Supplementary materials for this article are available online.
研究动机与目标
- 开发一种在特征数量p与样本量n相当或更大的情况下,针对广义线性模型(GLMs)中特征选择的稳健FDR控制框架。
- 解决现有方法(如Benjamini-Hochberg和knockoff过滤器)的局限性,这些方法依赖于渐近正态性,在高维情形下对缩放和偏差敏感。
- 提出一种无尺度方法,仅依赖于原假设下镜像统计量的对称性,从而提升有限样本下的性能。
- 在中等维(p/n → κ ∈ (0,1))和高维(p ≫ n)设置下,分别基于最大似然估计(MLE)和去偏Lasso,建立理论上的FDR控制。
提出的方法
- 通过数据分割或高斯镜像法获得两个渐近独立的真实系数估计,为每个特征构造镜像统计量。
- 利用原假设下镜像统计量抽样分布关于零对称的特性,实现FDR控制。
- 在中等维设置下,使用最大似然估计(MLE)估计系数,并基于分割数据或镜像数据上的成对MLE构造镜像统计量。
- 在高维设置下,使用去偏Lasso估计系数,并通过分割数据构造镜像统计量,以保持计算可行性。
- 应用数据相关的阈值,选择镜像统计量超过临界值的特征,确保在温和正则性和稀疏性条件下实现FDR控制。
- 利用无尺度特性:对镜像统计量进行任意常数缩放不会影响选择结果,从而增强有限样本下的稳健性。
实验结果
研究问题
- RQ1能否为GLMs开发一种不依赖检验统计量渐近正态性的无尺度FDR控制框架?
- RQ2如何构造镜像统计量以确保在原假设下具有渐近对称性,从而在高维GLM设置中实现稳健的FDR控制?
- RQ3所提出的方法在中等维和高维设置下是否能保持FDR控制,并且在统计功效上优于Benjamini-Hochberg和knockoff过滤器?
- RQ4该镜像统计量框架能否扩展至逻辑回归以外的GLMs(如负二项和泊松模型),并在一般协方差结构下适用?
- RQ5在p/n → κ ∈ (0,1)的GLMs中,数据分割法和高斯镜像法在何种理论条件下可实现渐近FDR控制?
主要发现
- 所提方法在中等维设置下(p/n → κ ∈ (0,1))使用基于MLE的镜像统计量,实现了渐近FDR控制,且无需知道经典MLE渐近理论中困扰的偏差和方差缩放因子。
- 在高维设置下(p ≫ n),该方法使用去偏Lasso估计构造镜像统计量,并在稀疏性和正则性条件下实现了FDR控制。
- 模拟结果表明,该方法在多种GLM族(包括逻辑、负二项和线性模型)中均能将FDR控制在名义水平(如q=0.1)附近。
- 该方法在统计功效上始终优于Benjamini-Hochberg程序和knockoff过滤器,尤其在有限样本和相关设计下表现更优。
- 实证结果表明,该方法对相关结构(恒定成对、部分和托普利茨相关)具有鲁棒性,在多样化设置下均保持稳定的FDR控制和高统计功效。
- 该方法的无尺度特性确保镜像统计量的常数缩放不会影响选择结果,从而在实际应用中提升可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。