Skip to main content
QUICK REVIEW

[Paper Review] A calibrated BISG for inferring race from surname and geolocation

Philip Greengard, Andrew Gelman|arXiv (Cornell University)|Apr 18, 2023
Names, Identity, and Discrimination Research4 citations
TL;DR

This paper introduces a raking-based calibration method that significantly improves Bayesian Improved Surname Geocoding (BISG) by correcting systematic biases caused by BISG’s flawed assumption of conditional independence between surname and geolocation given race. The method enhances accuracy and calibration on voter registration data, reducing undercounting of Hispanic and API populations by up to 25% and 15% respectively in predominantly White counties.

ABSTRACT

Bayesian Improved Surname Geocoding (BISG) is a ubiquitous tool for predicting race and ethnicity using an individual's geolocation and surname. Here we demonstrate that statistical dependence of surname and geolocation within racial/ethnic categories in the United States results in biases for minority subpopulations, and we introduce a raking-based improvement. Our method augments the data used by BISG--distributions of race by geolocation and race by surname--with the distribution of surname by geolocation obtained from state voter files. We validate our algorithm on state voter registration lists that contain self-identified race/ethnicity.

Motivation & Objective

  • To identify and correct systematic biases in BISG arising from its assumption that surname and geolocation are conditionally independent given race.
  • To improve the accuracy of race/ethnicity imputation for underrepresented subpopulations, especially in areas with high White population density.
  • To develop a simple, interpretable, and widely applicable calibration method that maintains BISG’s accessibility while enhancing performance.
  • To validate the improved method on real-world voter registration data from Florida and North Carolina, where race/ethnicity is labeled.
  • To demonstrate that BISG’s independence assumption leads to consistent undercounting of minority groups, undermining fairness in downstream applications.

Proposed method

  • Apply raking to adjust BISG predictions so that the estimated race/ethnicity distribution across geographies matches known, observed margins from labeled voter data.
  • Use the same labeled voter registration dataset for both training and validation to isolate the impact of BISG’s conditional independence assumption.
  • Calibrate BISG predictions by iteratively adjusting probabilities to match known population-level race/ethnicity distributions in each geographic unit.
  • Construct a fully self-consistent BISG model using exact census factors on the same data used for validation, eliminating distributional mismatch.
  • Compare the performance of raking-calibrated predictions against standard BISG using absolute and relative error metrics across counties.
  • Use subsampling and calibration maps to evaluate worst-case miscalibration and assess robustness across different subpopulations.

Experimental results

Research questions

  • RQ1Does BISG’s assumption of conditional independence between surname and geolocation given race lead to systematic biases in race/ethnicity estimation?
  • RQ2How does the undercounting of minority subpopulations by BISG vary across geographic regions with differing racial compositions?
  • RQ3Can a raking-based calibration method significantly improve BISG’s accuracy and calibration without requiring sensitive personal data?
  • RQ4To what extent do BISG’s errors in estimating subpopulation sizes affect downstream applications such as election law, healthcare, and lending?
  • RQ5Is the bias in BISG consistent across different states and demographic contexts, and can it be corrected using known population margins?

Key findings

  • In Florida, BISG underestimates Hispanic populations by 25% and API populations by 15% in counties where at least 85% of registered voters are non-Hispanic White.
  • The raking-based calibration method significantly reduces absolute and relative errors in subpopulation estimation compared to standard BISG across all counties in Florida and North Carolina.
  • The proposed method improves calibration by aligning predicted probabilities with known population margins, reducing worst-case miscalibration across predicted probability intervals.
  • BISG’s conditional independence assumption leads to consistent, systematic errors that are not random, particularly affecting small and minority subpopulations.
  • The raking-based approach maintains interpretability and does not require additional sensitive data, making it suitable for deployment in fairness-sensitive domains.
  • The study isolates the impact of BISG’s independence assumption by training and validating on the same labeled dataset, revealing previously unmeasured bias sources.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.