[论文解读] The Secret Lives of Names? Name Embeddings from Social Media
本文提出了一种基于推特数据训练的公开可访问姓名嵌入,实证表明其在性别、种族和国籍分类任务中的表现与私有的基于电子邮件的嵌入相当或更优。主要贡献在于发布了推特姓名嵌入(www.name-prism.com),并提供了证据表明这些嵌入通过编码传统人口统计特征未捕捉到的潜在社会经济与族群差异,能够提升寿命预测模型的性能。
Your name tells a lot about you: your gender, ethnicity and so on. It has been shown that name embeddings are more effective in representing names than traditional substring features. However, our previous name embedding model is trained on private email data and are not publicly accessible. In this paper, we explore learning name embeddings from public Twitter data. We argue that Twitter embeddings have two key advantages: extit{(i)} they can and will be publicly released to support research community. extit{(ii)} even with a smaller training corpus, Twitter embeddings achieve similar performances on multiple tasks comparing to email embeddings. As a test case to show the power of name embeddings, we investigate the modeling of lifespans. We find it interesting that adding name embeddings can further improve the performances of models using demographic features, which are traditionally used for lifespan modeling. Through residual analysis, we observe that fine-grained groups (potentially reflecting socioeconomic status) are the latent contributing factors encoded in name embeddings. These were previously hidden to demographic models, and may help to enhance the predictive power of a wide class of research studies.
研究动机与目标
- 使用推特数据开发公开可用的姓名嵌入,以克服基于私有电子邮件嵌入的局限性。
- 评估推特姓名嵌入在人口统计分类任务中是否能够达到或超越基于电子邮件的嵌入性能。
- 证明姓名嵌入在提升寿命预测模型性能方面具有预测能力,超越传统人口统计变量。
- 通过残差分析揭示姓名嵌入中编码的潜在社会经济与族群差异,这些差异未被标准人口统计变量捕捉。
提出的方法
- 在公开的推特数据上使用无监督技术训练姓名嵌入,利用沟通模式捕捉文化与社会同质性。
- 从推特互动中构建多种嵌入变体:提及、关注者、被关注者以及聚合嵌入,以探索不同的社交网络信号。
- 使用线性回归模型基于人口统计特征(出生年份、州、性别、种族、国籍)和姓名嵌入预测寿命。
- 对寿命预测模型进行残差分析,识别由姓名嵌入解释的未解释方差。
- 分析155对昵称/正式姓名及20个高频姓氏,检测对正式姓名或特定族群姓氏的偏好偏差。
- 应用统计假设检验(p < 0.01)以验证观察到的姓名形式与姓氏偏好偏差。
实验结果
研究问题
- RQ1基于推特的姓名嵌入是否能在人口统计分类任务中实现与私有电子邮件嵌入相当的性能?
- RQ2姓名嵌入是否能提升寿命预测模型的准确性,超越传统人口统计变量?
- RQ3姓名嵌入中编码了哪些未被标准人口统计变量捕捉到的潜在社会经济或族群差异?
- RQ4是否存在系统性偏差,使得正式姓名比昵称获得更高的寿命收益,且该偏差是否与社会经济地位相关?
- RQ5姓名嵌入是否能捕捉到细微的群体身份(例如阿什肯纳齐犹太人、斯堪的纳维亚人)并与其更长的寿命相关?
主要发现
- 尽管训练语料库较小,推特姓名嵌入在性别、种族和国籍分类任务中的表现与基于电子邮件的嵌入相当。
- 提及嵌入在性别识别任务中优于电子邮件嵌入,表明推特提及中存在更强的性别同质性。
- 聚合推特嵌入整体表现最佳,并已公开发布于 www.name-prism.com。
- 在所有32种测试配置中,引入姓名嵌入均显著提升了寿命预测模型性能,p值均小于0.01,表明具有统计显著性。
- 残差分析显示,姓名嵌入编码了细微的社会经济差异,例如对正式姓名和阿什肯纳齐犹太人姓氏的偏好,这些差异未被标准人口统计变量捕捉。
- 统计检验确认了对正式姓名的显著偏好(p < 0.01):在155对姓名中,使用电子邮件嵌入时74%的配对显示正式形式的寿命收益更高,使用推特嵌入时该比例为60%。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。