Skip to main content
QUICK REVIEW

[論文レビュー] Cluster-based name embeddings reduce ethnic disparities in record linkage quality under realistic name corruption: evidence from the North Carolina Voter Registry

Joseph Lam, Mario Cortina-Borja|arXiv (Cornell University)|Jan 12, 2026
Data Quality and Management被引用数 0
ひとこと要約

この論文は、クラスタベースの名寄せ埋め込みとTF補正姓比較を組み合わせることで、現実的な名前の破損下における民族別の過少結合の格差を、レコード連結の中で低減することを示している。NC Voter Registryデータを用いた検証。

ABSTRACT

Differential ethnic-based record linkage errors can bias epidemiologic estimates. Prior evidence often conflates heterogeneity in error mechanisms with unequal exposure to error. Using snapshots of the North Carolina Voter Registry (Oct 2011-Oct 2022), we derived empirical name-discrepancy profiles to parameterise realistic corruptions. From an Oct 2022 extract (n=848,566), we generated five replicate corrupted datasets under three settings that separately varied mechanism heterogeneity and exposure inequality, and linked records back to originals using unadjusted Jaro-Winkler, Term Frequency (TF)-adjusted Jaro-Winkler, and a cluster-based forename-embedding comparator combined with TF-adjusted surname comparison. We evaluated false match rate (FMR), missed match rate (MMR) and white-centric disparities. At a fixed MMR near 0.20, overall error rates and ethnic disparities diverged substantially by model under disproportionate exposure to corruption. Term-frequency (TF)-adjusted Jaro-Winkler achieved very low overall FMR (0.55% (95% CI 0.54-0.57)) at overall MMR 20.34% (20.30-20.39), but large white-centric under-linkage disparities persisted: Hispanic voters had 36.3% (36.1-36.6) and Non-Hispanic Black voters 8.6% (8.6-8.7) higher FMRs compared to Non-Hispanic White groups. Relative to unadjusted string similarity, TF adjustment reduced these disparities (Hispanic: +60.4% (60.1-60.7) to +36.3%; Black: +13.1% (13.0-13.2) to +8.6%). The cluster-based forename-embedding model reduced missed-match disparities further (Hispanic: +10.2% (9.8-10.3); Black: +0.6% (0.4-0.7)), but at a cost of increasing overall FMR (4.28% (4.22-4.35)) at the same threshold. Unequal exposure to identifier error drove substantially larger disparities than mechanism heterogeneity alone; cluster-based embeddings markedly narrowed under-linkage disparities beyond TF adjustment.

研究の動機と目的

  • 疫学的推定における民族別のレコード連結誤差を動機づけ、定量化する。
  • NC Voter Registryデータの経験的差異プロファイルから現実的な名前の破損をパラメータ化する。
  • 破損シナリオ全体で、未補正、TF補正、クラスタベース埋め込みの連結手法を比較する。
  • 誤りへの不均等な曝露が連結品質の格差に寄与する程度を評価する。

提案手法

  • NC Voter Registryのスナップショット(2011年10月〜2022年10月)から経験的名前差異プロファイルを構築する。
  • 2022年10月抽出データ長さ848,566から、機構の不均一性と曝露不平等の三設定下で五つの再現破損データセットを生成する。
  • 三つの比較子を用いてオリジナルとレコードをリンクする:未補正Jaro-Winkler、TF補正Jaro-Winkler、TF補正姓比較付きクラスタベースの名寄せ埋め込み。
  • 各手法での偽一致率(FMR)、見逃し一致率(MMR)、白人中心の格差を評価する。
  • 一定のMMR付近で固定して全体の誤差率と格差を比較する。

実験結果

リサーチクエスチョン

  • RQ1現実的な名前の破損と誤りの曝露が異なる場合、異なる名前連結アルゴリズムはどのように機能するか。
  • RQ2クラスタベースの名寄せ埋め込みは、文字列類似性だけと比べて偽一致・見逃しの格差を低減するか。
  • RQ3識別子誤りへの不均等な曝露が、機構の不均一性だけと比べて格差に与える影響はどの程度か。
  • RQ4TF補正の類似性指標とクラスタベース埋め込みは、全体の誤差トレードオフを考慮しつつ、白人中心の過少結合格差を狭めることができるか。

主な発見

  • TF補正Jaro-Winklerは全体のFMRを非常に低く抑える(約0.55%)一方、MMRが約20.34%のときでも白人中心の格差はなお大きい。
  • 未補正比較ではヒスパニック族および非ヒスパニック黒人のFMRは非ヒスパニック白人より高く、TF補正により格差は低減する。
  • クラスタベースの名寄せ埋め込みとTF補正姓比較を組み合わせると、見逃し一致の格差はさらに減少(ヒスパニック +10.2% vs +36.3%; 黒人 +0.6% vs +8.6%)するが、全体のFMRは4.28%へ増加する。
  • 同じ閾値で、誤りへの不均等な曝露は機構の不均一性だけより格差を大きく生み出し、クラスタベースの埋め込みはTF補正を超えて過少結合の格差を大幅に狭める。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。