Skip to main content
QUICK REVIEW

[论文解读] Freshman or Fresher? Quantifying the Geographic Variation of Internet Language

Vivek Kulkarni, Bryan Perozzi|arXiv (Cornell University)|Oct 22, 2015
Language and cultural evolution参考文献 39被引用 11
一句话总结

本文提出 geodist,一种新颖的统计方法,通过区域特定的词嵌入和零模型,检测并量化在线语言中具有地理意义的语义、句法和词汇变异,确保统计显著性。研究发现,由于全球化的影响,英式英语与美式英语的语义差异正在逐渐缩小。

ABSTRACT

We present a new computational technique to detect and analyze statistically significant geographic variation in language. Our meta-analysis approach captures statistical properties of word usage across geographical regions and uses statistical methods to identify significant changes specific to regions. While previous approaches have primarily focused on lexical variation between regions, our method identifies words that demonstrate semantic and syntactic variation as well. We extend recently developed techniques for neural language models to learn word representations which capture differing semantics across geographical regions. In order to quantify this variation and ensure robust detection of true regional differences, we formulate a null model to determine whether observed changes are statistically significant. Our method is the first such approach to explicitly account for random variation due to chance while detecting regional variation in word meaning. To validate our model, we study and analyze two different massive online data sets: millions of tweets from Twitter spanning not only four different countries but also fifty states, as well as millions of phrases contained in the Google Book Ngrams. Our analysis reveals interesting facets of language change at multiple scales of geographic resolution -- from neighboring states to distant continents. Finally, using our model, we propose a measure of semantic distance between languages. Our analysis of British and American English over a period of 100 years reveals that semantic variation between these dialects is shrinking.

研究动机与目标

  • 检测并量化词使用中超越单纯词汇差异的、具有统计显著性的地理变异。
  • 通过捕捉区域间的语义和句法变异,弥补现有方法的不足。
  • 开发一种稳健的方法,利用零模型区分真实的区域语言变化与随机变异。
  • 在大规模、多分辨率数据集(推文和 Google Ngrams)上评估该方法,覆盖多个国家和美国各州。
  • 提出一种新的无监督语义距离度量方法,用于分析方言随时间的趋同与分化。

提出的方法

  • 扩展神经语言模型,学习捕捉地理区域间语义差异的区域特定词嵌入。
  • 应用统计零模型,评估观察到的区域间词使用差异是否具有显著性,而非偶然所致。
  • 采用元分析方法,聚合区域间词使用统计特性,检测显著变化。
  • 利用在地理位置标记文本(如推文)和书籍 n-gram 上训练的词嵌入,大规模建模区域变异。
  • 提出一种基于不同区域词嵌入相似性的新方法,衡量方言间的语义距离。
  • 在多个数据集和地理分辨率上验证该方法,包括美国 50 个州和四个英语国家。

实验结果

研究问题

  • RQ1哪些词在地理区域间表现出超越词汇差异的、具有统计显著性的语义或句法变异?
  • RQ2如何在大规模在线文本中区分真实的区域语言变化与随机变异?
  • RQ3英式英语与美式英语方言在语义上随时间的趋同程度如何?
  • RQ4区域词用变异在社交媒体和出版文本中如何产生并传播?
  • RQ5我们能否以无监督、数据驱动的方式量化方言间的语义距离?

主要发现

  • 该方法成功检测到如 'test' 和 'schedule' 等词的语义变异:在印度 'test' 指板球比赛,而在其他地区指考试;在英国 'schedule' 指文本附件,而美国则无此用法。
  • 使用相同方法,可在多个尺度上检测到词用的地理变异,从相邻的美国州到遥远的大洲。
  • 过去 100 年间,英式英语与美式英语之间的语义距离持续减小,表明语义趋同趋势增强。
  • 零模型显著提升了检测准确性,通过过滤因随机抽样导致的虚假区域差异。
  • 该方法能识别社交媒体中的代码混用和区域特有用法,展现出对多样化语言现象的鲁棒性。
  • 所提出的语义距离度量表明,文化与技术全球化推动了方言趋同,尤其在英语中表现显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。