Skip to main content
QUICK REVIEW

[论文解读] Addressing Census data problems in race imputation via fully Bayesian Improved Surname Geocoding and name supplements

Kosuke Imai, Santiago Olivella|arXiv (Cornell University)|May 12, 2022
Data-Driven Disease Surveillance被引用 8
一句话总结

本文提出了一种全贝叶斯改进姓氏地理编码方法(fBISG),并利用美国南部选民档案中的姓名补充数据,以解决美国人口普查数据中的两个关键问题:少数族裔群体的零计数问题以及姓氏文件中缺少姓氏的问题。该方法显著提升了种族推断的准确性——尤其在亚裔和西班牙裔选民中表现突出,AUCROC最高提升15%,假阴性率降低18个百分点,同时在各族裔群体中保持了良好的校准性。

ABSTRACT

Prediction of individual's race and ethnicity plays an important role in social science and public health research. Examples include studies of racial disparity in health and voting. Recently, Bayesian Improved Surname Geocoding (BISG), which uses Bayes' rule to combine information from Census surname files with the geocoding of an individual's residence, has emerged as a leading methodology for this prediction task. Unfortunately, BISG suffers from two Census data problems that contribute to unsatisfactory predictive performance for minorities. First, the decennial Census often contains zero counts for minority racial groups in the Census blocks where some members of those groups reside. Second, because the Census surname files only include frequent names, many surnames -- especially those of minorities -- are missing from the list. To address the zero counts problem, we introduce a fully Bayesian Improved Surname Geocoding (fBISG) methodology that accounts for potential measurement error in Census counts by extending the naive Bayesian inference of the BISG methodology to full posterior inference. To address the missing surname problem, we supplement the Census surname data with additional data on last, first, and middle names taken from the voter files of six Southern states where self-reported race is available. Our empirical validation shows that the fBISG methodology and name supplements significantly improve the accuracy of race imputation across all racial groups, and especially for Asians. The proposed methodology, together with additional name data, is available via the open-source software WRU.

研究动机与目标

  • 解决由于人口普查数据限制导致贝叶斯改进姓氏地理编码(BISG)在预测少数族裔种族时表现不佳的问题。
  • 纠正少数族裔种族群体在人口普查区块中出现零计数的问题,该问题会削弱BISG的预测准确性。
  • 解决人口普查姓氏文件中对少数族裔群体姓氏覆盖不足的问题,特别是那些不常见姓氏的群体。
  • 提升所有主要种族类别中种族推断模型的预测准确性和校准性,尤其针对亚裔和西班牙裔群体。
  • 通过wru软件包发布增强方法的开源实现。

提出的方法

  • 提出一种全贝叶斯框架(fBISG),用完整的后验推断替代标准BISG中的朴素贝叶斯近似,以考虑人口普查计数中的测量误差。
  • 使用贝叶斯公式建模种族预测:P(R_i | S_i, G_i) ∝ P(S_i | R_i)P(R_i | G_i),其中S_i为姓氏,G_i为地理位置,R_i为种族。
  • 通过将区块级别的种族比例视为具有分层先验的随机变量,纳入人口普查区块层面种族比例的不确定性,从而实现完整的后验推断而非点估计。
  • 通过补充来自美国南部六个州选民档案中的额外姓名(包括名字和中间名),扩充标准姓氏-种族词典,以提升对代表性不足姓氏的覆盖度。
  • 采用分层建模结构,整合姓氏与地理区域之间的信息,提升对罕见姓氏和低计数区块的估计精度。
  • 使用来自同一南部州选民档案的自报种族数据,通过校准曲线对预测结果进行校准,并验证性能。

实验结果

研究问题

  • RQ1少数族裔群体在人口普查区块中出现零计数,如何影响标准BISG的预测表现?
  • RQ2由于人口普查姓氏文件中排除了不常见姓氏,对少数族裔群体的种族推断准确性会降低到何种程度?
  • RQ3一种考虑人口普查计数测量误差的全贝叶斯方法,能否提升种族预测的稳健性与准确性?
  • RQ4从选民档案中添加名字和中间名,能在多大程度上提升种族推断模型的准确性和覆盖度?
  • RQ5所提出的方法是否能维持或改善不同种族群体预测概率的校准性?

主要发现

  • 与标准BISG相比,fBISG方法使所有五个主要种族群体的平均AUCROC提升约7个百分点,其中亚裔群体的提升最高,达15个百分点(从0.82提升至0.94)。
  • 完整指定的模型使所有种族群体的分类准确率平均提升14%,其中亚裔群体的增益最为显著。
  • 在实施所有改进措施后,非白人选民的假阴性率平均降低18个百分点。
  • 白人选民的假阳性率降低近16个百分点,表明对白人与非白人个体的区分能力得到改善。
  • 当应用所有增强措施后,所有种族群体的预测概率均保持良好校准,AUCROC值均超过0.95。
  • 包含名字和中间名可提升准确性而不损害校准性,但当包含中间名时,西班牙裔和亚裔群体出现轻微校准性下降,可能与姓名使用中的文化差异有关。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。