[论文解读] Cluster-based name embeddings reduce ethnic disparities in record linkage quality under realistic name corruption: evidence from the North Carolina Voter Registry
The paper shows that cluster-based forename embeddings combined with TF-adjusted surname comparison reduces under-linkage disparities by ethnicity in record linkage under realistic name corruptions, using NC Voter Registry data.
Differential ethnic-based record linkage errors can bias epidemiologic estimates. Prior evidence often conflates heterogeneity in error mechanisms with unequal exposure to error. Using snapshots of the North Carolina Voter Registry (Oct 2011-Oct 2022), we derived empirical name-discrepancy profiles to parameterise realistic corruptions. From an Oct 2022 extract (n=848,566), we generated five replicate corrupted datasets under three settings that separately varied mechanism heterogeneity and exposure inequality, and linked records back to originals using unadjusted Jaro-Winkler, Term Frequency (TF)-adjusted Jaro-Winkler, and a cluster-based forename-embedding comparator combined with TF-adjusted surname comparison. We evaluated false match rate (FMR), missed match rate (MMR) and white-centric disparities. At a fixed MMR near 0.20, overall error rates and ethnic disparities diverged substantially by model under disproportionate exposure to corruption. Term-frequency (TF)-adjusted Jaro-Winkler achieved very low overall FMR (0.55% (95% CI 0.54-0.57)) at overall MMR 20.34% (20.30-20.39), but large white-centric under-linkage disparities persisted: Hispanic voters had 36.3% (36.1-36.6) and Non-Hispanic Black voters 8.6% (8.6-8.7) higher FMRs compared to Non-Hispanic White groups. Relative to unadjusted string similarity, TF adjustment reduced these disparities (Hispanic: +60.4% (60.1-60.7) to +36.3%; Black: +13.1% (13.0-13.2) to +8.6%). The cluster-based forename-embedding model reduced missed-match disparities further (Hispanic: +10.2% (9.8-10.3); Black: +0.6% (0.4-0.7)), but at a cost of increasing overall FMR (4.28% (4.22-4.35)) at the same threshold. Unequal exposure to identifier error drove substantially larger disparities than mechanism heterogeneity alone; cluster-based embeddings markedly narrowed under-linkage disparities beyond TF adjustment.
研究动机与目标
- Motivate and quantify differential ethnic-based record linkage errors in epidemiologic estimates.
- Parameterize realistic name corruptions using empirical discrepancy profiles from NC Voter Registry data.
- Compare unadjusted, TF-adjusted, and cluster-based embedding linkage methods across corruption scenarios.
- Assess how unequal exposure to errors contributes to disparities in linkage quality.
提出的方法
- Construct empirical name-discrepancy profiles from NC Voter Registry snapshots (Oct 2011-Oct 2022).
- Generate five replicate corrupted datasets from an Oct 2022 extract (n=848,566) under three settings of mechanism heterogeneity and exposure inequality.
- Link records to originals using three comparators: unadjusted Jaro-Winkler, TF-adjusted Jaro-Winkler, and cluster-based forename-embedding with TF-adjusted surname comparison.
- Evaluate false match rate (FMR), missed match rate (MMR), and white-centric disparities across methods.
- Analyze performance at fixed MMR near 0.20 to compare overall error rates and disparities.
实验结果
研究问题
- RQ1不同名字链接算法在现实名字损坏且暴露于错误程度不同的情况下表现如何?
- RQ2基于簇的名字符嵌入是否在仅字符串相似度的基础上减少白人以外族裔在误匹配和漏匹配上的差异?
- RQ3不平等暴露于标识符错误对差异与机制异质性有何影响?
- RQ4在考虑总体错误权衡的同时,TF 调整的相似性度量和基于簇的嵌入能否缩窄白人主导的低关联差异?
主要发现
- TF 调整的 Jaro-Winkler 在 MMR ~20.34% 时实现非常低的总体 FMR (~0.55%),但白人主导差异仍然显著。
- 在未调整的比较下,西班牙裔和非西班牙裔黑人群体的FMR高于非西班牙裔白人群体;通过 TF 调整可减小差异。
- 基于名字符嵌入的簇聚与 TF 调整的姓氏比较进一步减小漏匹配差异(西班牙裔 +10.2% 对 +36.3%;黑人 +0.6% 对 +8.6%),但把总体 FMR 提升到 4.28%。
- 在同一阈值下,不平等暴露于损坏比机制异质性本身驱动更大差异,且基于簇的嵌入在超越 TF 调整的情况下显著缩小低关联差异。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。