[论文解读] Identification, Interpretability, and Bayesian Word Embeddings
本文提出贝叶斯词嵌入(Bayesian Word Embeddings)结合自动相关性确定(Automatic Relevance Determination),以解决标准词嵌入在社会科学研究中存在的一致性识别与可解释性问题。通过将嵌入建模为贝叶斯潜在变量,并利用理论驱动的词语锚定维度,该方法生成可直接用于回归分析、具备可解释性的词表示。研究揭示了美国总统就职演说中1945年后国际主义修辞的显著下降趋势,以及外交文件中精英好战情绪与美国敌对性外交政策行动增加之间的显著关联。
Social scientists have recently turned to analyzing text using tools from natural language processing like word embeddings to measure concepts like ideology, bias, and affinity. However, word embeddings are difficult to use in the regression framework familiar to social scientists: embeddings are are neither identified, nor directly interpretable. I offer two advances on standard embedding models to remedy these problems. First, I develop Bayesian Word Embeddings with Automatic Relevance Determination priors, relaxing the assumption that all embedding dimensions have equal weight. Second, I apply work identifying latent variable models to anchor the dimensions of the resulting embeddings, identifying them, and making them interpretable and usable in a regression. I then apply this model and anchoring approach to two cases, the shift in internationalist rhetoric in the American presidents' inaugural addresses, and the relationship between bellicosity in American foreign policy decision-makers' deliberations. I find that inaugural addresses became less internationalist after 1945, which goes against the conventional wisdom, and that an increase in bellicosity is associated with an increase in hostile actions by the United States, showing that elite deliberations are not cheap talk, and helping confirm the validity of the model.
研究动机与目标
- 解决标准词嵌入在识别与可解释性方面的不足,以克服其在社会科学研究中常见回归模型应用的障碍。
- 开发一种基于贝叶斯框架的词嵌入方法,通过自动相关性确定(ARD)先验实现对各维度的差异化正则化。
- 利用理论驱动的词语锚定嵌入维度,以实现因果推断所需的识别与可解释性。
- 将该方法应用于两个真实语料:美国总统就职演说和解密的《美国外交关系文件》(declassified Foreign Relations of the United States, FRUS)外交文件。
- 通过展示模型测度与外部冲突数据的相关性,并揭示新颖的历史趋势,对模型进行验证。
提出的方法
- 将词嵌入形式化为贝叶斯潜在变量模型,采用变分贝叶斯推断估计后验分布。
- 在嵌入维度上引入自动相关性确定(ARD)先验,以实现不同维度的差异化加权,并识别无关维度。
- 应用理想点建模中的识别技术(如 Rivers, 2003;Clinton et al., 2004),利用语义上有意义的词语锚定嵌入维度。
- 采用两步锚定过程:首先,识别高对比度的词对(如“和平”与“战争”)以定义维度端点;其次,沿这些维度对所有词语进行尺度化。
- 使用泊松广义线性模型(Poisson generalized linear model)检验外交文件中嵌入的敌意程度与实际冲突事件之间的因果关系。
- 通过对比模型输出与外部冲突事件统计量,评估模型拟合度与稳健性,实现模型验证。
实验结果
研究问题
- RQ1自1945年以来,美国总统就职演说中对国际主义的修辞强调是否有所下降?这一趋势是否与国际关系领域的传统认知相悖?
- RQ2精英外交讨论中的敌意程度在多大程度上能够预测美国实际的敌对性外交政策行为?
- RQ3能否利用具有ARD先验和维度锚定的贝叶斯词嵌入,构建政治文本中语义概念的可解释、可回归测度?
- RQ4嵌入的修辞测度是否与独立的冲突行为数据集相关,从而支持模型的建构效度?
- RQ5如何使词嵌入具备可识别性与可解释性,以支持因果社会科学研究中的推断?
主要发现
- 对就职演说的分析显示,1945年后国际主义修辞显著下降,与传统预期中全球参与度持续上升的观点相悖。
- 贝叶斯词嵌入模型成功利用理论驱动的词语锚点识别并解释了嵌入维度,使其具备回归分析的适用性。
- 前两个星期的敌意程度得分每提高一个标准差,美国发起的敌对事件数量显著增加,该结果通过95%置信区间的泊松广义线性模型(Poisson GLM)得到验证。
- 该模型的敌意程度量表与外部冲突数据高度相关,支持其作为精英修辞测量工具的有效性与可靠性。
- 该方法能够检测到传统文档级“文本即数据”方法无法识别的语义演变,如国际主义的衰落。
- 使用ARD先验可识别并降低无关嵌入维度的权重,从而减少噪声,提升模型的可解释性与效率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。