[论文解读] Discursive Landscapes and Unsupervised Topic Modeling in IR: A Validation of Text-As-Data Approaches through a New Corpus of UN Security Council Speeches on Afghanistan
本文通过分析2001–2017年联合国安全理事会关于阿富汗的演讲新语料库,验证了国际关系领域中无监督主题建模(LDA)的有效性。采用混合方法——将LDA主题与定性文献进行比较,并分析发言者-主题网络——结果表明,LDA能够可靠地捕捉话语格局,并揭示出如‘妇女与人权’等有意义的联盟,从而增强了文本作为数据方法在国际关系研究中的可信度。
The recent turn towards quantitative text-as-data approaches in IR brought new ways to study the discursive landscape of world politics. Here seen as complementary to qualitative approaches, quantitative assessments have the advantage of being able to order and make comprehensible vast amounts of text. However, the validity of unsupervised methods applied to the types of text available in large quantities needs to be established before they can speak to other studies relying on text and discourse as data. In this paper, we introduce a new text corpus of United Nations Security Council (UNSC) speeches on Afghanistan between 2001 and 2017; we study this corpus through unsupervised topic modeling (LDA) with the central aim to validate the topic categories that the LDA identifies; and we discuss the added value, and complementarity, of quantitative text-as-data approaches. We set-up two tests using mixed- method approaches. Firstly, we evaluate the identified topics by assessing whether they conform with previous qualitative work on the development of the situation in Afghanistan. Secondly, we use network analysis to study the underlying social structures of what we will call 'speaker-topic relations' to see whether they correspondent to know divisions and coalitions in the UNSC. In both cases we find that the unsupervised LDA indeed provides valid and valuable outputs. In addition, the mixed-method approaches themselves reveal interesting patterns deserving future qualitative research. Amongst these are the coalition and dynamics around the 'women and human rights' topic as part of the UNSC debates on Afghanistan.
研究动机与目标
- 验证无监督主题建模(LDA)作为分析国际关系中大规模话语的可靠方法。
- 开发并发布一个关于2001–2017年联合国安全理事会关于阿富汗的演讲的全新、全面语料库,供学术研究使用。
- 评估LDA识别出的主题是否与既有关于阿富汗冲突演变的定性研究发现一致。
- 通过网络分析探索发言者-主题关系的社会结构,以识别联合国安理会在实际中的联盟与分化。
- 展示定量文本作为数据方法与定性话语分析在国际关系研究中的互补性。
提出的方法
- 构建了一个包含1,052篇2001至2017年联合国安全理事会关于阿富汗的演讲的新标注语料库。
- 应用潜在狄利克雷分布(LDA)进行无监督主题建模,以识别演讲中的潜在主题结构。
- 通过将LDA主题与既有关于阿富汗政治与安全发展问题的定性研究发现进行对比,开展混合方法验证。
- 通过将每篇演讲与其主导的LDA主题关联,绘制‘发言者-主题关系’图谱,并构建该关系网络。
- 使用网络分析方法考察发言者-主题关联中的结构性模式,识别出聚类与核心人物。
- 通过定性合理性检验及与已知地缘政治动态的一致性,评估主题的连贯性与可解释性。
实验结果
研究问题
- RQ1LDA在联合国安全理事会关于阿富汗的演讲中识别出的主题,是否与文献中关于该冲突的经验性与定性确立的主题相一致?
- RQ2从LDA中推导出的发言者-主题关系,在多大程度上反映了联合国安理会在实际中的政治联盟与分化?
- RQ3联合国安理会关于阿富汗的演讲中的主题动态如何随时间演变?其演变是否与冲突的总体发展轨迹一致?
- RQ4无监督主题建模能否有效捕捉复杂冲突背景下多边外交的话语格局?
- RQ5通过主题建模与发言者-主题模式的网络分析交叉分析,能揭示出关于外交协调与议题重要性的哪些新见解?
主要发现
- LDA生成的主题,如‘妇女与人权’和‘安全部门改革’,与关于阿富汗冲突演变的定性研究发现高度吻合。
- ‘妇女与人权’主题成为共识焦点,获得多个常任与非常任理事国跨区域的持续支持。
- 对发言者-主题关系的网络分析揭示了明显的合作集群,包括围绕人权与安全问题的核心联盟。
- LDA模型成功识别出10个稳定且可解释的主题,主题连贯性高,主题间具有明确的语义区分。
- 本研究证明,通过混合方法验证后,无监督主题建模能够产生有效、可解释且具有政策相关性的输出结果。
- 将LDA与网络分析相结合,揭示了动态外交模式,如联盟的演变与议题特定联盟的形成,这些模式值得进一步开展定性研究。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。