Skip to main content
QUICK REVIEW

[論文レビュー] Where in the World are You? Geolocation and Language Identification in Twitter

Mark Graham, Scott A. Hale|Oxford University Research Archive (ORA) (University of Oxford)|Aug 3, 2013
Multilingual Education and Policy参考文献 4被引用数 11
ひとこと要約

本稿は、Twitterにおける自動的地理特定および言語識別手法の評価を行い、3つの言語検出ツールをTwitterのUI言語設定と人間による検証済みラベルと比較することで実施している。ユーザーが入力したプロフィール上の場所は、実際の物理的場所の代替として不適切であることが判明し、言語識別ツールの正確性には顕著な差が見られる。これは、ソーシャルメディア研究において、慎重な方法論的選択が求められることを示唆している。

ABSTRACT

The movements of ideas and content between locations and languages are unquestionably crucial concerns to researchers of the information age, and Twitter has emerged as a central, global platform on which hundreds of millions of people share knowledge and information. A variety of research has attempted to harvest locational and linguistic metadata from tweets in order to understand important questions related to the 300 million tweets that flow through the platform each day. However, much of this work is carried out with only limited understandings of how best to work with the spatial and linguistic contexts in which the information was produced. Furthermore, standard, well-accepted practices have yet to emerge. As such, this paper studies the reliability of key methods used to determine language and location of content in Twitter. It compares three automated language identification packages to Twitter's user interface language setting and to a human coding of languages in order to identify common sources of disagreement. The paper also demonstrates that in many cases user-entered profile locations differ from the physical locations users are actually tweeting from. As such, these open-ended, user-generated, profile locations cannot be used as useful proxies for the physical locations from which information is published to Twitter.

研究の動機と目的

  • Twitterデータにおける自動言語識別ツールの信頼性を評価すること。
  • ユーザーが入力したプロフィール上の場所が、実際の物理的場所をどれだけ的確に反映しているかを評価すること。
  • 自動言語検出手法を、TwitterのUI言語設定と人間による検証済みラベルと比較すること。
  • Twitterのコンテンツにおける言語的メタデータと実際の地理的場所の間に生じる不一致を特定すること。
  • 研究者がTwitterデータを用いた地理的特定および言語識別において、最良の実践法を支援する根拠に基づいた指針を提供すること。

提案手法

  • 人間による検証済み言語ラベルと比較して、3つの自動言語識別パッケージ(例:langdetect、fastText、その他のツール)を評価した。
  • 言語識別におけるベンチマークとして、TwitterのUI言語設定を用いた。
  • ユーザーが入力したプロフィール上の場所を収集・分析し、ツイートからの実際の地理的位置データと比較した。
  • ツイート内容の人間によるコード化を実施し、真の言語ラベルを確立した。
  • 自動化ツール、UI言語、人間によるラベルの間の一致度を測る統計的分析を実施した。
  • 地理座標を用いて、プロフィール上の場所と実際のツイートの場所との間の空間的不一致を評価した。

実験結果

リサーチクエスチョン

  • RQ1人間による検証済みラベルと比較した場合、自動言語識別ツールはTwitter上でどれほど正確か?
  • RQ2ユーザーが入力したTwitterプロフィール上の場所は、ツイートが発信された実際の地理的場所をどの程度的確に反映しているか?
  • RQ3TwitterのUI言語設定は、ツイート内容の実際の言語とどの程度一致しているか?
  • RQ4自動言語検出ツールと人間による検証済みラベルとの間に生じる不一致の原因は何か?
  • RQ5ユーザーが入力したプロフィール上の場所は、Twitterデータ分析において物理的地理的位置の代替として信頼できるか?

主な発見

  • Twitter上のユーザーが入力したプロフィール上の場所は、しばしば不正確であり、ツイートが発信された物理的場所を的確に反映していない。
  • 自動言語識別ツールの性能には顕著な差が見られ、一部のツールは同じデータセットで90%以上の正確性を示す一方、他のツールは80%未満にとどまる。
  • TwitterのUI言語設定は、ツイート内容の言語と一貫して一致せず、一部のケースでは誤分類を引き起こしている。
  • プロフィール上の場所と実際のツイートの場所との間に生じる不一致は広範にわたり、多くのユーザーが自身の地理的ルートとは一致しない場所を記載している。
  • 人間による検証済み言語ラベルは、自動化ツールやTwitterのUI言語設定よりも信頼性の高いベンチマークである。
  • 本研究では、あらゆる言語および地理的文脈において、1つの自動化ツールが常に他のツールを上回るとは限らないことが判明し、文脈に応じた手法選択の重要性が強調された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。