Skip to main content
QUICK REVIEW

[论文解读] Valid Post-Selection Inference in High-Dimensional Approximately Sparse Quantile Regression Models

Alexandre Belloni, Victor Chernozhukov|arXiv (Cornell University)|Dec 27, 2013
Statistical Methods and Inference被引用 4
一句话总结

该论文针对近似稀疏条件下的高维分位数回归模型,提出了有效的后模型选择推断方法,利用正交得分函数以确保对模型选择误差的鲁棒性。即使在解释变量数量超过样本量的情况下,该方法仍能建立回归系数的统一渐近正态性及有效的置信区域,且无需满足beta-min条件。

ABSTRACT

This work proposes new inference methods for a regression coefficient of interest in a (heterogeneous) quantile regression model. We consider a high-dimensional model where the number of regressors potentially exceeds the sample size but a subset of them suffice to construct a reasonable approximation to the conditional quantile function. The proposed methods are (explicitly or implicitly) based on orthogonal score functions that protect against moderate model selection mistakes, which are often inevitable in the approximately sparse model considered in the present paper. We establish the uniform validity of the proposed confidence regions for the quantile regression coefficient. Importantly, these methods directly apply to more than one variable and a continuum of quantile indices. In addition, the performance of the proposed methods is illustrated through Monte-Carlo experiments and an empirical example, dealing with risk factors in childhood malnutrition.

研究动机与目标

  • 开发在p ≫ n的高维分位数回归模型中回归系数的推断方法。
  • 确保在高维设定下模型选择误差普遍存在时,置信区域的 validity。
  • 将推断方法扩展至多维处理变量及分位数索引的连续体。
  • 通过利用正交得分函数,消除对beta-min条件的需求。
  • 在弱正则性和稀疏性假设下,提供统一有效的置信区域。

提出的方法

  • 使用对混淆函数gτ的一阶估计误差具有鲁棒性的正交得分函数。
  • 采用ℓ1-惩罚分位数回归和后模型选择估计进行初始模型拟合。
  • 对密度加权方程应用异方差后Lasso,以偏回归出混淆变量。
  • 构建两种估计量:一种通过Neyman型得分统计量,另一种通过在选中变量上进行密度加权分位数回归。
  • 为估计量实施一个枢轴线性表示,以实现有效的推断。
  • 利用Neyman型得分统计量的渐近卡方分布(自由度为1)进行置信带构造。

实验结果

研究问题

  • RQ1在高维近似稀疏模型中,是否可以对模型选择后的分位数回归系数构造有效的置信区域?
  • RQ2当高维设定下模型选择误差不可避免时,如何保持推断的有效性?
  • RQ3所提出的方法能否处理多维处理变量及分位数索引的连续体?
  • RQ4该推断程序是否对混淆函数gτ的非正则估计具有鲁棒性?
  • RQ5该方法是否避免了渐近有效性所需的beta-min条件?

主要发现

  • 所提出的估计量具有根n一致性和渐近正态性,且具有枢轴线性表示。
  • 基于估计标准误的置信区域在弱矩条件和稀疏性条件下,渐近覆盖概率为1−ξ。
  • 在原假设下,Neyman型得分统计量渐近服从自由度为1的卡方分布,从而可实现有效的置信带。
  • 在阵列渐近框架下,该方法对广泛的数据生成过程类具有统一有效性。
  • 置信区域无需将回归系数与零分离(即无需beta-min条件)即可保持有效性。
  • 蒙特卡洛实验和对儿童营养不良风险因素的实证应用验证了该方法在小样本下的表现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。