Skip to main content
QUICK REVIEW

[论文解读] Dialectograms: Machine Learning Differences between Discursive Communities

Thyge Enggaard, August Lohse|arXiv (Cornell University)|Feb 11, 2023
Social Media and Politics被引用 5
一句话总结

本文提出了一种名为方言图(dialectograms)的新方法,这是一种无监督方法,利用词嵌入技术对话语社群在焦点词使用上的差异进行可视化和量化分析。与仅关注词级别差异的现有方法不同,该方法通过分析整个嵌入空间,揭示了细微的语言差异,例如情感极化和政治评估的分歧。该方法在两个美国政治子版块(r/Political_Webcomic 和 r/Political_StackExchange)中得到应用,成功克服了现有测量方法对罕见词或多义词的偏好所带来的局限性。

ABSTRACT

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word embeddings are complex, high-dimensional spaces and a focus on identifying differences only captures a fraction of their richness. Here, we take a step towards leveraging the richness of the full embedding space, by using word embeddings to map out how words are used differently. Specifically, we describe the construction of dialectograms, an unsupervised way to visually explore the characteristic ways in which each community use a focal word. Based on these dialectograms, we provide a new measure of the degree to which words are used differently that overcomes the tendency for existing measures to pick out low frequent or polysemous words. We apply our methods to explore the discourses of two US political subreddits and show how our methods identify stark affective polarisation of politicians and political entities, differences in the assessment of proper political action as well as disagreement about whether certain issues require political intervention at all.

研究动机与目标

  • 解决现有方法仅关注话语社群中孤立词语差异的局限性。
  • 利用词嵌入空间的全部丰富性,而非依赖稀疏的高维向量比较。
  • 开发一种可视化和定量工具,用于探索社群如何以独特而典型的方式使用特定词语。
  • 创建一种稳健的词汇分歧度量方法,避免对低频词或多义词产生偏差。
  • 将该方法应用于现实世界的政治话语,揭示不同社群间政治语言的结构性差异。

提出的方法

  • 使用预训练语言模型为多个话语社群的文本语料构建词嵌入。
  • 针对每个焦点词,提取其嵌入向量,并分析其在社群特异性嵌入空间中的分布。
  • 通过可视化词语在各社群语义空间内的空间聚类和方向偏移,生成方言图。
  • 使用嵌入空间分歧的统计度量(如余弦相似度和方向方差)来量化使用模式的差异。
  • 应用一种新的分歧度量方法,同时考虑语义中心性和分布范围,降低对罕见词或歧义词的敏感性。
  • 在两个美国政治子版块(r/Political_Webcomic 和 r/Political_StackExchange)上验证该方法,以分析政治话语的差异。

实验结果

研究问题

  • RQ1如何在超越孤立词语比较的基础上,可视化并量化话语社群之间词语使用差异的全貌?
  • RQ2词嵌入的空间结构在捕捉社群特异性语言模式方面起到什么作用?
  • RQ3为何现有词汇分歧度量方法因对低频词或多义词的偏好,而无法捕捉有意义的差异?
  • RQ4政治子版块在政治术语的情感与评价性使用上,表现出多大程度的分歧?
  • RQ5方言图能否揭示政治话语中的结构性差异,例如对政治干预必要性的分歧?

主要发现

  • 方言图成功揭示了在对政治人物和政治实体的描绘中,两个子版块之间存在显著的情感极化,各自表现出截然不同的情感极性。
  • 该方法识别出社群在评估政治行为时存在系统性差异,例如对某些行为是否被视为合法或非法的认知不同。
  • 社群在对特定议题是否需要政治干预的感知上存在差异,这反映在相关术语的语义定位上。
  • 所提出的分歧度量方法通过降低对低频词和多义词的敏感性,优于现有指标,从而产生更稳定、更易解释的结果。
  • 分析表明,词嵌入中蕴含着丰富且结构化的使用差异,这些差异无法通过传统基于词级别的比较方法捕捉。
  • 该方法揭示了传统自然语言处理技术仅关注孤立词语差异时所无法察觉的、细致且社群特异的语义模式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。