[論文レビュー] Detecting the Age of Twitter Users.
本稿では、名前や言語的特徴に依存せずに、誰をフォローしているかというソーシャルネットワーク構造に基づいて、Twitterユーザーの年齢を言語に依存しない方法で検出する手法を提案する。7億件のアカウントにベイジアン分類モデルを適用した結果、F1スコア0.86を達成し、ユーザーが提供する生年月日データや言語固有のモデルを必要とせずに、正確な年齢推定が可能である。
Twitter provides an extremely rich and open source of data for studying human behaviour at scale. It has been used to advance our understanding of social network structure, the viral flow of information and how new ideas develop. Enriching Twitter with demographic information would permit more precise science and better generalisation to the real world. The only demographic indicators associated with a Twitter account are the free text name, location and description fields. We show how the age of most Twitter accounts can be inferred with high accuracy using the structure of the social graph. Besides classical social science applications, there are obvious privacy and child protection implications to this discovery. Previous work on Twitter age detection has focussed on either user-name or linguistic features of tweets. A shortcoming of the user-name approach is that it requires real names (Twitter names are often false) and census data from each user's (unknown) birth country. Problems with linguistic approaches are that most Twitter users do not tweet (the median number of Tweets is 4) and a different model must be learnt for each language. To address these issues, we devise a language-independent methodology for determining the age of Twitter users from data that is native to the Twitter ecosystem. Roughly 150,000 Twitter users specify an age in their free text description field. We generalize this to the entire Twitter network by showing that age can be predicted based on what or whom they follow. We adopt a Bayesian classification paradigm, which offers a consistent framework for handling uncertainty in our data, e.g., inaccurate age descriptions or spurious edges in the graph. Working within this paradigm we have successfully applied age detection to 700 million Twitter accounts with an F1 Score of 0.86.
研究の動機と目的
- ユーザーが提供する名前や言語的コンテンツに依存せず、Twitterにおける信頼できる人口統計データの欠如に対処し、正確な年齢検出を可能にすること。
- 実名(しばしば偽物)や言語固有のモデルに依存する従来のアプローチの限界を克服し、多言語対応やツイートが少ないユーザーに対しても実用的であるようにすること。
- 多様な集団にわたる、主にソーシャルグラフ関係(フォロー関係)に基づく、ネイティブなTwitterデータから年齢を推定するスケーラブルで汎用性の高い手法を開発すること。
- ベイジアン推論を用いて、年齢記述の不確実性やノイズの多いソーシャルグラフエッジを一貫して扱うフレームワークを提供すること。
- より正確な社会科学の研究や、子供の保護メカニズムの向上を可能にするために、大規模にユーザーの年齢を正確に推定できるようにすること。
提案手法
- ユーザーがフォローするユーザーの集合、つまりソーシャルグラフを主な特徴として用い、ユーザー名やツイート内容に依存しない年齢予測を実現する。
- 年齢記述の不確実性やソーシャルネットワークにおける誤った接続(スパイアスエッジ)をモデル化するベイジアン分類フレームワークを適用し、頑健な推論を可能にする。
- プロフィール記述に自ら年齢を記載している約15万件のユーザーをシードセットとして用い、広範なネットワーク全体にわたる年齢パターンを一般化する。
- 年齢関連の同質的結合(homophily)の仮定に基づき、ユーザーの社交圈权内のユーザーの年齢分布を確率的モデルで用いて、年齢グループを推定する。
- 大規模グラフに適した効率的な推論技術を用いて、7億件のTwitterアカウントにスケーリングする。
- 多クラス分類アプローチを採用し、年齢グループ(例:18–24、25–34など)を用いることで、多様な集団にわたる精度と一般化性能のバランスを取る。
実験結果
リサーチクエスチョン
- RQ1ユーザーの年齢は、言語的特徴や名前に基づかずに、ソーシャルネットワーク構造のみを用いて正確に予測可能か?
- RQ2ユーザーが誰をフォローしているかに依存する年齢検出の性能は、異なる年齢層でどのように変化するか?
- RQ3プロフィールに自ら年齢を記載しているユーザーのデータを、Twitter全体のユーザーに一般化するモデルを学習するためにどれほど有効に使えるか?
- RQ4ベイジアンフレームワークは、年齢ラベルの不確実性やノイズの多いソーシャルグラフデータをどのように処理するか?
- RQ5この手法は、言語固有のチューニングなしに、数百万〜数十億人のユーザーにスケーリングして適用可能か?
主な発見
- 本手法は、7億件のTwitterアカウントに対して年齢グループを予測する際、F1スコア0.86を達成し、大規模なスケールでも高い精度を示した。
- 言語に依存しないため、言語固有のモデルやツイートの言語的分析を必要としない。
- ユーザー名やツイート内容に依存する従来の手法よりも優れており、特にツイートが少なく(中央値4件)、言語的特徴が乏しいユーザーに対しても優位性を示した。
- ベイジアンフレームワークは、自ら報告した年齢の不確実性や誤ったソーシャル接続を効果的に処理し、耐障害性を向上させた。
- 15万人のユーザーに自ら年齢を報告している小さなシードセットから、Twitter全体のネットワークにうまく一般化された。
- 本研究では、言語的・名前的手がかりがなくとも、ソーシャルグラフ構造そのものが、信頼できる年齢推定に十分な信号を含んでいることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。