[论文解读] Common change point estimation in panel data from the least squares and maximum likelihood viewpoints
本文通过最小二乘法和最大似然(ML)方法,建立了面板数据中常见变点估计量的收敛速率和渐近分布,表明ML方法需要更强的信噪比可识别性条件,但其估计量具有更低的渐近方差。本文提出了一种数据驱动的自适应程序,可在不事先知晓渐近 regime 的情况下构造有效的置信区间。
We establish the convergence rates and asymptotic distributions of the common break change-point estimators, obtained by least squares and maximum likelihood in panel data models and compare their asymptotic variances. Our model assumptions accommodate a variety of commonly encountered probability distributions and, in particular, models of particular interest in econometrics beyond the commonly analyzed Gaussian model, including the zero-inflated Poisson model for count data, and the probit and tobit models. We also provide novel results for time dependent data in the signal-plus-noise model, with emphasis on a wide array of noise processes, including Gaussian process, MA$(\infty)$ and $m$-dependent processes. The obtained results show that maximum likelihood estimation requires a stronger signal-to-noise model identifiability condition compared to its least squares counterpart. Finally, since there are three different asymptotic regimes that depend on the behavior of the norm difference of the model parameters before and after the change point, which cannot be realistically assumed to be known, we develop a novel data driven adaptive procedure that provides valid confidence intervals for the common break, without requiring a priori knowledge of the asymptotic regime the problem falls in.
研究动机与目标
- 比较面板数据模型中基于最小二乘法和最大似然法的共同变点估计量的渐近性质。
- 研究在两种估计准则下实现一致性和渐近正态性所需的可识别性条件。
- 将分析扩展至非高斯分布,包括零膨胀泊松、probit 和 Tobit 模型等非高斯分布。
- 开发一种数据驱动的自适应推断程序,可在不事先知晓渐近 regime 的情况下,为所有三种渐近 regime 构造有效的变点置信区间。
- 分析误差结构中的时间依赖性影响,包括 MA(∞) 和 m-依赖过程,而不仅限于 i.i.d. 噪声。
提出的方法
- 在一般模型假设下,推导了最小二乘法和 ML 估计量在共同变点上的收敛速率和渐近分布,涵盖指数族和非高斯分布。
- 引入两种信噪比条件:LS 的 SNR1 和 ML 的 SNR2,表明 SNR2 更强,说明 ML 对可识别性要求更严格。
- 采用归一化目标函数方法,推导变点估计量的极限分布,其在变点两侧分别对应不同的极限过程(布朗运动)。
- 在适当的矩条件和混合条件之下,应用 Lyaounov 中心极限定理,证明了重标化目标函数弱收敛于布朗运动的泛函。
- 提出一种基于数据驱动选择渐近 regime 的自适应程序,利用对误差方差参数的经验估计来校准置信区间。
- 采用条件概率测度(P**)处理高概率事件中矩和方差界成立的情形,确保在条件测度下收敛于分布。
实验结果
研究问题
- RQ1在渐近分布的方差和可识别性条件方面,面板数据中共同变点的最小二乘法与最大似然估计量的比较如何?
- RQ2在最小二乘法与最大似然估计下,实现一致性和渐近正态性所需的最小信噪比条件是什么?
- RQ3能否构建一个单一的自适应推断程序,使其在不事先知晓 regime 的情况下,为所有三种渐近 regime 构造有效的变点置信区间?
- RQ4当误差过程具有时间依赖性(如 MA(∞)、m-依赖)而非 i.i.d. 噪声时,估计量的渐近性质如何变化?
- RQ5在零膨胀泊松或 Tobit 等非高斯模型中,最大似然估计是否能检测到最小二乘法无法识别的结构变化?
主要发现
- 最大似然估计量所需的信噪比条件(SNR2)强于最小二乘估计量(SNR1),表明在弱信号条件下 ML 更为苛刻。
- 尽管可识别性要求更强,ML 估计量的渐近方差小于最小二乘估计量,因此在模型正确设定时更为高效。
- 对于正态分布数据,两种方法下的一致性信噪比条件相同,且渐近分布一致。
- 本文确立了变点估计量的三种不同渐近 regime,取决于变点前后参数的范数差,而该差值事先未知。
- 提出了一种自适应推断程序,利用经验方差估计选择正确的渐近 regime,并在不事先知晓 regime 的情况下构造有效的置信区间。
- 证明了变点估计量的极限分布收敛于布朗运动的泛函,其在变点左右两侧具有不同的方差参数(γ_L 和 γ_R),反映了估计中的不对称性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。