[论文解读] Confidence intervals for multiple isotonic regression and other monotone models
本文通过利用块最大-最小和最小-最大估计量,为多重等渗回归及其他单调模型构建了渐近精确的置信区间。通过结合这些估计量的顶点信息,消除了极限分布中的干扰参数,得到一个不依赖未知导数的枢轴渐近分布,从而在正则条件下实现了具有精确覆盖概率的Oracle长度置信区间。
We consider the problem of constructing pointwise confidence intervals in the multiple isotonic regression model. Recently, [HZ19] obtained a pointwise limit distribution theory for the so-called block max-min and min-max estimators [FLN17] in this model, but inference remains a difficult problem due to the nuisance parameter in the limit distribution that involves multiple unknown partial derivatives of the true regression function. In this paper, we show that this difficult nuisance parameter can be effectively eliminated by taking advantage of information beyond point estimates in the block max-min and min-max estimators. Formally, let $\hat{u}(x_0)$ (resp. $\hat{v}(x_0)$) be the maximizing lower-left (resp. minimizing upper-right) vertex in the block max-min (resp. min-max) estimator, and $\hat{f}_n$ be the average of the block max-min and min-max estimators. If all (first-order) partial derivatives of $f_0$ are non-vanishing at $x_0$, then the following pivotal limit distribution theory holds: $$ \sqrt{n_{\hat{u},\hat{v}}(x_0)}\big(\hat{f}_n(x_0)-f_0(x_0)\big) ightsquigarrow σ\cdot \mathbb{L}_{1_d}. $$ Here $n_{\hat{u},\hat{v}}(x_0)$ is the number of design points in the block $[\hat{u}(x_0),\hat{v}(x_0)]$, $σ$ is the standard deviation of the errors, and $\mathbb{L}_{1_d}$ is a universal limit distribution free of nuisance parameters. This immediately yields confidence intervals for $f_0(x_0)$ with asymptotically exact confidence level and oracle length. Notably, the construction of the confidence intervals, even new in the univariate setting, requires no more efforts than performing an isotonic regression for once using the block max-min and min-max estimators, and can be easily adapted to other common monotone models. Extensive simulations are carried out to support our theory.
研究动机与目标
- 为解决多重等渗回归中极限分布涉及未知偏导数时构建有效置信区间的挑战。
- 通过利用块最大-最小和最小-最大估计量提供的点估计之外的额外信息,消除渐近分布中的干扰参数。
- 建立一个通用且不依赖未知导数的枢轴极限分布,实现精确推断。
- 将该方法扩展至其他单调模型,包括单调密度估计和面板计数模型。
提出的方法
- 定义块最大-最小和最小-最大估计量,其中 $\widehat{u}(x_0)$ 和 $\widehat{v}(x_0)$ 分别为块 $[\widehat{u}(x_0), \widehat{v}(x_0)]$ 中使下左顶点最大化的点和使上右顶点最小化的点。
- 构造平均估计量 $\widehat{f}_n(x_0) = \frac{1}{2}(\widehat{f}_n^-(x_0) + \widehat{f}_n^+(x_0))$ 以提高估计稳定性。
- 将块大小 $n_{\widehat{u},\widehat{v}}(x_0)$ 用作随机缩放因子,对估计误差 $\sqrt{n_{\widehat{u},\widehat{v}}(x_0)}(\widehat{f}_n(x_0) - f_0(x_0))$ 进行标准化。
- 证明在 $x_0$ 处偏导数非零的条件下,标准化误差依分布收敛于 $\sigma \cdot \mathbb{L}_{\mathbf{1}_d}$,即一个与干扰参数无关的通用极限分布。
- 建立极限分布 $\mathbb{L}_{\mathbf{1}_d}$ 为分布自由且枢轴的,从而可实现具有渐近精确覆盖概率和最小长度的置信区间。
- 将该方法推广至其他单调模型,包括单调密度估计、当前状态数据、面板计数模型以及具有形状约束的广义线性模型。
实验结果
研究问题
- RQ1能否在多重等渗回归中构建具有渐近精确覆盖概率和Oracle长度的置信区间?
- RQ2如何消除极限分布中涉及未知偏导数的干扰参数?
- RQ3枢轴极限分布理论能否超越等渗回归,推广至其他单调模型?
- RQ4块最大-最小和最小-最大估计量结构是否足以消除对未知导数的依赖?
主要发现
- 标准化估计误差 $\sqrt{n_{\widehat{u},\widehat{v}}(x_0)}(\widehat{f}_n(x_0) - f_0(x_0))$ 依分布收敛于 $\sigma \cdot \mathbb{L}_{\mathbf{1}_d}$,即一个不包含干扰参数的通用极限分布。
- 极限分布 $\mathbb{L}_{\mathbf{1}_d}$ 为枢轴且分布自由,可实现具有渐近精确覆盖概率和最小长度的置信区间。
- 该方法除一次使用块估计量执行等渗回归外,无需额外计算,计算效率高。
- 该方法可推广至其他单调模型,包括单调密度估计、当前状态数据、面板计数模型以及具有单调性约束的广义线性模型。
- 模拟结果支持理论发现,证实了在有限样本中具有准确的覆盖概率和Oracle长度表现。
- 关键洞见在于,块估计量的顶点信息可消除未知偏导数,从而解决了形状约束模型推断中的长期难题。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。