[论文解读] Small area estimation of general finite-population parameters based on grouped data
本文提出了一种基于模型的小面积估计新方法,适用于使用分组数据(如收入等级频率)的一般有限总体参数。该方法采用潜在变量方法,结合线性混合模型与多项分布似然函数,将组概率与辅助变量关联,通过吉布斯抽样和带重要性抽样的马尔可夫链蒙特卡洛期望最大化算法实现经验贝叶斯估计。
This paper proposes a new model-based approach to small area estimation of general finite-population parameters based on grouped data or frequency data, which is often available from sample surveys. Grouped data contains information on frequencies of some pre-specified groups in each area, for example the numbers of households in the income classes, and thus provides more detailed insight about small areas than area-level aggregated data. A direct application of the widely used small area methods, such as the Fay-Herriot model for area-level data and nested error regression model for unit-level data, is not appropriate since they are not designed for grouped data. The newly proposed method adopts the multinomial likelihood function for the grouped data. In order to connect the group probabilities of the multinomial likelihood and the auxiliary variables within the framework of small area estimation, we introduce the unobserved unit-level quantities of interest which follows the linear mixed model with the random intercepts and dispersions after some transformation. Then the probabilities that a unit belongs to the groups can be derived and are used to construct the likelihood function for the grouped data given the random effects. The unknown model parameters (hyperparameters) are estimated by a newly developed Monte Carlo EM algorithm using an efficient importance sampling. The empirical best predicts (empirical Bayes estimates) of small area parameters can be calculated by a simple Gibbs sampling algorithm. The numerical performance of the proposed method is illustrated based on the model-based and design-based simulations. In the application to the city level grouped income data of Japan, we complete the patchy maps of the Gini coefficient as well as mean income across the country.
研究动机与目标
- 解决在仅有分组数据(如收入等级频率)可用时,缺乏对一般有限总体参数的可靠小面积估计器的问题。
- 克服现有Fay–Herriot模型和嵌套误差模型的局限性,这些模型因缺乏个体水平信息而不适用于分组数据。
- 构建一个统一框架,通过潜在个体水平变量和随机效应,将分组数据频率与辅助变量关联。
- 仅使用频率数据即可实现对复杂参数(如基尼系数和均值)的小面积水平估计。
- 提供一种计算上可行的估计程序,结合马尔可夫链蒙特卡洛期望最大化算法与吉布斯抽样,实现经验贝叶斯预测。
提出的方法
- 基于预定义组别的观测频率,使用多项分布似然函数对分组数据建模。
- 引入未观测到的潜在个体水平变量,代表感兴趣的真正值,其取值被限制在组区间内。
- 假设潜在变量服从具有随机截距和异方差误差的线性混合模型,以与辅助变量关联。
- 推导在潜在变量、随机效应和方差分量条件下的分组数据的似然函数。
- 使用带高效重要性抽样的马尔可夫链蒙特卡洛期望最大化算法,估计超参数(如方差分量)。
- 通过从潜在变量和随机效应的全条件分布中进行吉布斯抽样,计算经验最佳预测。
实验结果
研究问题
- RQ1当仅有分组数据(如收入等级频率)可用时,能否为一般有限总体参数开发基于模型的小面积估计框架?
- RQ2在混合模型框架下,如何利用潜在个体水平变量将分组数据频率与辅助变量关联?
- RQ3在具有分组数据和潜在变量的模型中,估计超参数的高效计算方法是什么?
- RQ4所提出的方法在小面积水平上估计复杂参数(如基尼系数和平均收入)时表现如何?
- RQ5当样本量较小时,直接估计器不可靠,该方法是否仍能产生可靠且稳定的小面积估计?
主要发现
- 所提方法成功实现了仅基于分组数据对一般有限总体参数(包括基尼系数和平均收入)的小面积估计。
- 带重要性抽样的马尔可夫链蒙特卡洛期望最大化算法即使在高维潜在变量空间中,也能提供稳定且准确的超参数估计。
- 通过吉布斯抽样可高效获得小面积参数的经验贝叶斯估计,且全条件分布可解析推导。
- 在模拟研究中,该方法在均方误差方面优于直接估计器,尤其在样本量有限的小面积中表现更优。
- 在日本1,265个市镇的收入数据上的应用,生成了基尼系数和平均收入的不规则分布图,展示了其实际应用价值。
- 该方法通过分层混合模型结构,有效处理了分组数据中的不确定性,实现了跨区域的信息借用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。