Skip to main content
QUICK REVIEW

[论文解读] Where in the World are You? Geolocation and Language Identification in Twitter

Mark Graham, Scott A. Hale|Oxford University Research Archive (ORA) (University of Oxford)|Aug 3, 2013
Multilingual Education and Policy参考文献 4被引用 11
一句话总结

本文通过将三种语言检测工具与Twitter的界面语言设置及人工验证标签进行对比,评估了在Twitter上自动地理定位和语言识别方法的效果。研究发现,用户提供的个人资料位置无法准确反映其实际物理位置,且语言识别工具的准确性差异显著,凸显了在社交媒体研究中需谨慎选择方法论。

ABSTRACT

The movements of ideas and content between locations and languages are unquestionably crucial concerns to researchers of the information age, and Twitter has emerged as a central, global platform on which hundreds of millions of people share knowledge and information. A variety of research has attempted to harvest locational and linguistic metadata from tweets in order to understand important questions related to the 300 million tweets that flow through the platform each day. However, much of this work is carried out with only limited understandings of how best to work with the spatial and linguistic contexts in which the information was produced. Furthermore, standard, well-accepted practices have yet to emerge. As such, this paper studies the reliability of key methods used to determine language and location of content in Twitter. It compares three automated language identification packages to Twitter's user interface language setting and to a human coding of languages in order to identify common sources of disagreement. The paper also demonstrates that in many cases user-entered profile locations differ from the physical locations users are actually tweeting from. As such, these open-ended, user-generated, profile locations cannot be used as useful proxies for the physical locations from which information is published to Twitter.

研究动机与目标

  • 评估Twitter数据中自动化语言识别工具的可靠性。
  • 评估用户输入的个人资料位置作为实际物理位置代理的准确性。
  • 将自动化语言检测方法与Twitter的界面语言设置及人工验证标签进行对比。
  • 识别Twitter内容中语言元数据与实际地理位置之间的差异。
  • 为研究人员在使用Twitter数据进行地理定位与语言识别时提供基于证据的最佳实践指导。

提出的方法

  • 将三种自动化语言识别软件包(如langdetect、fastText及其他)与人工验证的语言标签进行对比。
  • 使用Twitter的界面语言设置作为语言识别的基准。
  • 收集并分析用户提供的个人资料位置,并与推文的实际地理坐标数据进行对比。
  • 通过人工编码推文内容以建立真实语言标签的基准。
  • 开展统计分析,测量自动化工具、界面语言与人工标签之间的一致性。
  • 利用地理坐标评估个人资料位置与实际推文位置之间的空间错位情况。

实验结果

研究问题

  • RQ1与人工验证的语言标签相比,自动化语言识别工具在Twitter上的准确性如何?
  • RQ2Twitter个人资料中用户提供的位置在多大程度上反映了推文实际发布时的地理来源?
  • RQ3Twitter的界面语言设置与推文内容的实际语言在多大程度上保持一致?
  • RQ4自动化语言检测工具与人工验证标签之间不一致的来源是什么?
  • RQ5用户提供的个人资料位置能否在Twitter数据分析中可靠地用作物理地理定位的代理?

主要发现

  • Twitter用户提供的个人资料位置通常不准确,无法可靠反映推文实际发送的物理位置。
  • 自动化语言识别工具在性能上存在显著差异,部分工具在相同数据集上的准确率超过90%,而另一些则低于80%。
  • Twitter的界面语言设置与推文内容的实际语言并不总是一致,导致部分情况下出现误分类。
  • 个人资料位置与实际推文位置之间的错位现象普遍存在,许多用户列出的位置与其地理原籍不匹配。
  • 人工验证的语言标签比自动化工具或Twitter的界面语言设置更具可靠性。
  • 本研究揭示,没有任何一种自动化工具在所有语言和地理背景下始终优于其他工具,凸显了根据上下文选择方法的重要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。