Skip to main content
QUICK REVIEW

[论文解读] Inference for Rank-Rank Regressions

Denis Chetverikov, Daniel Wilhelm|arXiv (Cornell University)|Oct 24, 2023
Intergenerational and Educational Inequality Studies被引用 5
一句话总结

本文为秩-秩回归中的OLS推断提出了一个一致的渐近方差估计量,表明标准同方差和稳健方差估计量由于依赖于收入分布的依存结构(copula)而存在不一致性。作者推导了秩基回归的一般渐近理论,将其扩展至聚类、混合秩次及秩次回归变量模型,并证明错误的方差估计会导致置信区间过宽或过窄,从而改变对代际流动性的结论。

ABSTRACT

The slope coefficient in a rank-rank regression is a popular measure of intergenerational mobility. In this article, we first show that commonly used inference methods for this slope parameter are invalid. Second, when the underlying distribution is not continuous, the OLS estimator and its asymptotic distribution may be highly sensitive to how ties in the ranks are handled. Motivated by these findings we develop a new asymptotic theory for the OLS estimator in a general class of rank-rank regression specifications without imposing any assumptions about the continuity of the underlying distribution. We then extend the asymptotic theory to other regressions involving ranks that have been used in empirical work. Finally, we apply our new inference methods to two empirical studies on intergenerational mobility, highlighting the practical implications of our theoretical findings.

研究动机与目标

  • 识别并修正秩-秩回归中标准误估计量(同方差与Eicker-White)的不一致性,这些估计量依赖于收入分布的依存结构(copula)。
  • 为不假设连续分布或无协变量的秩-秩回归中的OLS估计量建立一般渐近理论。
  • 将渐近理论扩展至聚类秩-秩回归、以秩次结果与非秩次回归变量为特征的回归,以及以非秩次结果与秩次回归变量为特征的回归。
  • 为所有这些模型提出一致的方差估计量,以支持代际流动性实证研究中的有效推断。
  • 通过实证证明,错误的标准误会导致置信区间显著过长或过短,从而影响对不同地区或国家间流动性差异的结论。

提出的方法

  • 在一般条件下(包括非连续分布和协变量存在)推导秩-秩回归中OLS估计量的渐近分布,涵盖非连续分布和协变量的情形。
  • 证明同方差与Eicker-White方差估计量的概率极限会因秩次变量之间依存结构(copula)的不同而偏离真实渐近方差。
  • 通过考虑秩次中的依赖结构与结 ties 处理机制,提出一致的渐近方差估计量。
  • 将理论扩展至三种额外回归类型:聚类秩-秩回归(在共同尺度上计算秩次)、以秩次结果与非秩次回归变量为特征的回归,以及以非秩次结果与秩次回归变量为特征的回归。
  • 在R包(csranks)中实现这些估计量,并开发了用于实证应用的Stata命令。
  • 校准模拟实验,并将方法应用于三个真实世界数据集:美国代际流动性、实验室实验的复制研究,以及区域间流动性比较。
Figure 1: Variances achieving minimal and maximal differences within different families of copulas. Specifications and details can be found in Appendix A . “correct” refers to $\sigma^{2}$ , “hom” to $\sigma_{hom}^{2}$ , “EW” to $\sigma_{EW}^{2}$ , and “rho” to the rank correlation of the copula ach
Figure 1: Variances achieving minimal and maximal differences within different families of copulas. Specifications and details can be found in Appendix A . “correct” refers to $\sigma^{2}$ , “hom” to $\sigma_{hom}^{2}$ , “EW” to $\sigma_{EW}^{2}$ , and “rho” to the rank correlation of the copula ach

实验结果

研究问题

  • RQ1为何标准同方差与稳健方差估计量在秩-秩回归中无法一致地估计渐近方差?
  • RQ2父母与子女收入分布之间依存结构(copula)的形状如何影响常用方差估计量的偏误?
  • RQ3当分布非连续或包含协变量时,秩-秩回归中OLS估计量的正确渐近分布是什么?
  • RQ4在实证应用中,基于正确方差估计量的置信区间与基于标准估计量的置信区间有何不同?
  • RQ5这些推断差异在多大程度上影响了对美国通勤区之间或国家之间代际流动性结论的判断?

主要发现

  • 同方差与Eicker-White方差估计量的概率极限可能因收入分布的依存结构(copula)而显著过大或过小,导致标准推断失效。
  • 对于1940–1980年父子代际队列,同方差标准误比正确标准误最大高出30%,导致置信区间过于保守。
  • 在18项实验室实验的复制研究中,对于某一关键指标,同方差标准误约为正确标准误的2.5倍,使置信区间宽度翻倍。
  • Eicker-White估计量在多个案例中也产生了过长的置信区间,部分估计值比正确值大50%以上。
  • 当结 ties 处理方式不同时(如ω=0),部分同方差与Eicker-White置信区间显著短于正确区间,尤其在二值回归变量情况下更为明显。
  • 基于Spearman等级相关(假设边缘分布连续)的置信区间与秩-秩回归斜率的估计值中心位置不同,且当数据中存在点质量(pointmasses)时无效。
Figure 2: Spearman’s rank correlation (“Spearman”) and the slope coefficient from a rank-rank regression (“rank-rank reg”), where $Y$ is a child’s income and $X$ their parent’s income. The graphs in the first column (“bottom censoring”) show the two estimates when all parental incomes up to the $q$
Figure 2: Spearman’s rank correlation (“Spearman”) and the slope coefficient from a rank-rank regression (“rank-rank reg”), where $Y$ is a child’s income and $X$ their parent’s income. The graphs in the first column (“bottom censoring”) show the two estimates when all parental incomes up to the $q$

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。