[论文解读] Gibrat's law for cities: uniformly most powerful unbiased test of the Pareto against the lognormal
本文通过将一致最功效无偏检验(UMPUT)应用于2000年人口普查数据,解决了关于美国城市规模是否遵循齐夫定律(帕累托分布,指数为1)或对数正态分布的长期争议。研究发现,规模最大的1,000个城市遵循尾指数为1.4 ± 0.1的帕累托分布,在90%置信水平下拒绝齐夫定律;而较小城市则更符合对数正态分布,从而调和了以往相互矛盾的研究结果。
We address the general problem of testing a power law distribution versus a log-normal distribution in statistical data. This general problem is illustrated on the distribution of the 2000 US census of city sizes. We provide definitive results to close the debate between Eeckhout (2004, 2009) and Levy (2009) on the validity of Zipf's law, which is the special Pareto law with tail exponent 1, to describe the tail of the distribution of U.S. city sizes. Because the origin of the disagreement between Eeckhout and Levy stems from the limited power of their tests, we perform the {\em uniformly most powerful unbiased test} for the null hypothesis of the Pareto distribution against the lognormal. The $p$-value and Hill's estimator as a function of city size lower threshold confirm indubitably that the size distribution of the 1000 largest cities or so, which include more than half of the total U.S. population, is Pareto, but we rule out that the tail exponent, estimated to be $1.4 \pm 0.1$, is equal to 1. For larger ranks, the $p$-value becomes very small and Hill's estimator decays systematically with decreasing ranks, qualifying the lognormal distribution as the better model for the set of smaller cities. These two results reconcile the opposite views of Eeckhout (2004, 2009) and Levy (2009). We explain how Gibrat's law of proportional growth underpins both the Pareto and lognormal distributions and stress the key ingredient at the origin of their difference in standard stochastic growth models of cities \cite{Gabaix99,Eeckhout2004}.
研究动机与目标
- 解决Eeckhout(2004)与Levy(2007)之间长期存在的争议:美国城市规模是遵循帕累托分布(齐夫定律)还是对数正态分布。
- 解决以往检验方法统计功效有限的问题,特别是检测帕累托与对数正态分布尾部差异的能力不足。
- 应用一致最功效无偏检验(UMPUT),以明确评估帕累托分布原假设与对数正态分布备择假设。
- 通过严格的统计推断,判断帕累托分布的尾指数是否等于1(即齐夫定律是否成立),或显著不同。
提出的方法
- 作者应用一致最功效无偏检验(UMPUT),用于检验帕累托分布原假设与对数正态分布备择假设。
- 采用鞍点近似方法计算不同城市规模下限阈值下的UMPUT p值,确保在小样本情况下仍具有高精度。
- 使用Hill估计量来估计帕累托尾指数的倒数(α⁻¹),并基于UMPUT结果构建置信带,以评估统计显著性。
- 通过蒙特卡洛模拟验证UMPUT结果,确认p值与估计量行为的稳健性。
- 将该方法应用于美国2000年人口普查数据,聚焦于按规模从大到小排序的城市,通过调整下限阈值来评估尾部分布行为。
- 将实证结果与参数μ=7.28、σ=1.25的对数正态分布模拟数据进行比较,该参数与Eeckhout的模型相匹配。
实验结果
研究问题
- RQ1美国最大城市的规模分布更符合帕累托分布还是对数正态分布?其统计证据是什么?
- RQ2美国城市规模帕累托分布的尾指数是否等于1(即齐夫定律是否成立),或是否存在显著差异?
- RQ3为何以往研究使用Lilliefors检验(L检验)及其他方法,对城市规模分布形式得出了相互矛盾的结论?
- RQ4帕累托分布与对数正态分布的渐近尾部分布行为有何不同?能否通过实证数据加以区分?
- RQ5Gibrat的成比例增长定律在生成帕累托与对数正态分布中起什么作用?要产生尾指数≠1的帕累托尾部,还需哪些额外假设?
主要发现
- 一致最功效无偏检验(UMPUT)证实,规模最大的1,000个美国城市的规模分布为帕累托分布,且在较低阈值下p值保持较高,表明该模型在该范围内得到强有力支持。
- 尾指数α估计为1.4 ± 0.1,与1存在显著差异,因此在90%置信水平下拒绝齐夫定律。
- 当城市排名超过1,000名后,p值急剧下降,表明对数正态分布对较小城市的拟合更优。
- Hill估计量对尾指数倒数(α⁻¹)在前1,000个城市中保持约0.7的稳定水平,证实存在稳定的帕累托尾部;但对于更高排名(更大城市数),其值系统性下降,表明偏离帕累托行为。
- 参数μ=7.28、σ=1.25的模拟对数正态数据未表现出Hill估计量的平台现象,证实美国实证数据在上尾部分的行为与对数正态分布不一致。
- 本研究调和了Eeckhout(2004)的对数正态结论与Levy(2007)的帕累托结论,表明帕累托模型适用于最大城市,而对数正态模型更适用于整体及较小城市。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。