[论文解读] Nonparametric Estimation of the Regression Function in an Errors-in-Variables Model
本文提出了一种在错误变量模型中对回归函数进行非参数估计的方法,通过利用去卷积技术与自适应惩罚对比估计器。在误差密度和回归函数具有不同光滑性假设的条件下,建立了该估计器的最优收敛速率,证明当误差为普通光滑或超高光滑时,且回归函数与密度共享正则性特征时,该估计器可达到极小极大收敛速率。
We consider the regression model with errors-in-variables where we observe $n$ i.i.d. copies of $(Y,Z)$ satisfying $Y=f(X)+ξ, Z=X+σε$, involving independent and unobserved random variables $X,ξ,ε$. The density $g$ of $X$ is unknown, whereas the density of $σε$ is completely known. Using the observations $(Y\_i, Z\_i)$, $i=1,...,n$, we propose an estimator of the regression function $f$, built as the ratio of two penalized minimum contrast estimators of $\ell=fg$ and $g$, without any prior knowledge on their smoothness. We prove that its $\mathbb{L}\_2$-risk on a compact set is bounded by the sum of the two $\mathbb{L}\_2(\mathbb{R})$-risks of the estimators of $\ell$ and $g$, and give the rate of convergence of such estimators for various smoothness classes for $\ell$ and $g$, when the errors $ε$ are either ordinary smooth or super smooth. The resulting rate is optimal in a minimax sense in all cases where lower bounds are available.
研究动机与目标
- 为在协变量存在测量误差的错误变量模型中开发回归函数的非参数估计器。
- 解决当误差密度未知且可能为普通光滑或超高光滑时,估计回归函数的挑战。
- 在误差密度和基础函数的各种光滑性假设下,推导估计器的最优收敛速率。
- 证明当误差密度与回归函数具有匹配的正则性特征时,所提出的估计器可达到极小极大收敛速率。
- 研究当协变量密度同时被估计时估计器的性能,并确定估计器达到最优性的条件。
提出的方法
- 该方法采用去卷积方法来估计回归函数与协变量密度的乘积,记为 ℓ = f·g。
- 对 ℓ 和 g 使用惩罚对比估计器,维度选择基于最小化期望平方误差。
- 通过截断阶数 D_m̆_ℓ 和 D_m̆_g 的基展开构造估计器,以平衡偏差与方差。
- 应用塔拉格朗德型集中不等式来控制风险界中的经验过程项。
- 通过将经验过程分解为两部分(一部分涉及回归函数,另一部分涉及误差),推导出惩罚估计器的风险上界。
- 最终的 f 估计器通过估计的 ℓ 和 g 的比值得到,并附有其收敛速率的理论保证。
实验结果
研究问题
- RQ1当误差密度为普通光滑或超高光滑时,在错误变量模型中非参数回归的最优收敛速率是什么?
- RQ2在误差密度和回归函数的光滑性假设不同的情况下,所提出的估计器能否达到极小极大收敛速率?
- RQ3估计器的性能如何依赖于误差密度和协变量密度的光滑性?
- RQ4当回归函数比协变量密度更光滑时,估计器在何种条件下达到最优?
- RQ5该方法能否在未知光滑性参数的情况下自适应调整,而无需事先知道基础函数的正则性?
主要发现
- 当误差密度为普通光滑且 r_ℓ = r_g = 0 时,最优收敛速率为 O(n^{-2a^*/(2a^* + 2α + 1)}),其中 a^* = min(a_ℓ, a_g)。
- 当误差密度为超高光滑且 r_ℓ = r_g = 0 时,收敛速率为 O((ln n)^{-2a^*/ρ}),其快于任何负幂次的 n,但慢于任何 ln n 的幂次。
- 对于具有 r_ℓ > 0 和 r_g > 0 的超高光滑误差,收敛速率为 O(ln(n)^{(2α+1)/r*}/n),其中 r^* = min(r_ℓ, r_g),且该速率快于任何 ln n 的幂次。
- 当回归函数与协变量密度具有相同光滑性时,估计器达到极小极大收敛速率,如 Fan (1991a)、Butucea (2004) 和 Butucea 与 Tsybakov (2004) 所示。
- 当密度 g 比回归函数 f 更光滑时,估计器仍保持最优,但当 f 比 g 更光滑时,最优性未被建立,这凸显了基于比值的估计器的局限性。
- 最终估计器 ̂f_̆m_ℓ,̆m_g 的风险界由 ℓ 和 g 的估计器中最差风险主导,且在函数与误差分布的弱矩和正则性条件下该界成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。