Skip to main content
QUICK REVIEW

[論文レビュー] Demographics in Social Media Data for Public Health Research: Does it matter?

Nina Cesare, Christan Grant|arXiv (Cornell University)|Oct 30, 2017
Data-Driven Disease Surveillance参考文献 49被引用数 15
ひとこと要約

本稿では、公衆衛生研究におけるソーシャルメディアデータのデモグラフィック表現のギャップを是正するため、Twitterのユーザーネームから性別を推定するスケーラブルなアンサンブル分類器を提案する。7,953人のラベル付きユーザーで訓練された3つの分類器を組み合わせることで、82.8%の正確性を達成し、アルファヌメリック文字列のユーザーネームを有するユーザーの100%をカバーする。データマッチング手法に比べて正確性は低いが、カバレッジが優れているため、米国の食中毒傾向のより包括的な分析を可能にする。

ABSTRACT

Social media data provides propitious opportunities for public health research. However, studies suggest that disparities may exist in the representation of certain populations (e.g., people of lower socioeconomic status). To quantify and address these disparities in population representation, we need demographic information, which is usually missing from most social media platforms. Here, we propose an ensemble approach for inferring demographics from social media data. Several methods have been proposed for inferring demographic attributes such as, age, gender and race/ethnicity. However, most of these methods require large volumes of data, which makes their application to large scale studies challenging. We develop a scalable approach that relies only on user names to predict gender. We develop three separate classifiers trained on data containing the gender labels of 7,953 Twitter users from Kaggle.com. Next, we combine predictions from the individual classifiers using a stacked generalization technique and apply the ensemble classifier to a dataset of 36,085 geotagged foodborne illness related tweets from the United States. Our ensemble approach achieves an accuracy, precision, recall, and F1 score of 0.828, 0.851, 0.852 and 0.837, respectively, higher than the individual machine learning approaches. The ensemble classifier also covers any user with an alphanumeric name, while the data matching approach, which achieves an accuracy of 0.917, only covers 67% of users. Application of our method to reports of foodborne illness in the United States highlights disparities in tweeting by gender and shows that counties with a high volume of foodborne-illness related tweets are heavily overrepresented by female Twitter users.

研究の動機と目的

  • 公衆衛生研究に用いられるソーシャルメディアデータにおけるデモグラフィック表現の不均衡を是正すること。
  • デモグラフィックデータが欠落している状況において、Twitterのユーザーネームから性別を推定するスケーラブルな手法を開発すること。
  • ソーシャルメディアを用いた大規模な公衆衛生研究における性別予測のカバレッジと正確性を向上させること。
  • 食中毒に関するツイートを用いた公衆衛生監視において、デモグラフィックの不均衡が及ぼす影響を評価すること。

提案手法

  • 7,953人のTwitterユーザーの既知の性別ラベルを用いて訓練された3つの個別分類器の予測を統合するため、スタックド一般化を用いたアンサンブル分類器を構築する。
  • 分類器は入力特徴としてユーザーネームのみに依存しており、アルファヌメリック文字列のユーザーネームを有するユーザーに広範にカバー可能なスケーラビリティを実現する。
  • メタラーナーを用いてベースモデルの予測を統合することで、個々のモデルを上回る全体的なパフォーマンスを向上させる。
  • 米国の地理的にタグ付けされた食中毒関連ツイート36,085件のデータセットにこの手法を適用し、性別の表現状況を評価する。
  • 91.7%の正確性を達成するが、ユーザーの67%しかカバーしないデータマッチング手法と比較する。
  • 標準的な指標(正確性、適合率、再現率、F1スコア)を用いてアンサンブルモデルを評価する。

実験結果

リサーチクエスチョン

  • RQ1ユーザーネームのみを用いることで、公衆衛生研究におけるソーシャルメディアデータの性別推定が正確かつスケーラブルに行えるか?
  • RQ2外部データマッチングに依存する場合と比較して、ユーザーネームに依存する性別予測手法のカバレッジはどの程度か?
  • RQ3ソーシャルメディアデータにおけるデモグラフィックの不均衡が、公衆衛生監視の結果にどの程度影響を及ぼすか?
  • RQ4アンサンブル学習アプローチが、ユーザーネームからの個別分類器を上回る性能を示せるか?
  • RQ5ソーシャルメディアでの性別による報告の不均衡が、食中毒監視に及ぼす影響は何か?

主な発見

  • アンサンブル分類器はテストセットで82.8%の正確性、85.1%の適合率、85.2%の再現率、83.7%のF1スコアを達成した。
  • アンサンブル手法はアルファヌメリック文字列のユーザーネームを有するユーザーの100%をカバーしており、データマッチング手法(67%のカバレッジ)を著しく上回った。
  • データマッチング手法は高い正確性(91.7%)を達成したが、カバレッジが低く、正確性と包括性のトレードオフが顕著に現れた。
  • 食中毒ツイートの分析から、報告件数の多い郡では女性のTwitterユーザーが顕著に多く占めていることが明らかになった。
  • 本研究では、ソーシャルメディアデータにおけるデモグラフィックの不均衡が、公衆衛生監視の結果に顕著に影響を及ぼす可能性があることを示した。
  • 提案手法により、より代表的でスケーラブルなデモグラフィック推定が可能となり、より公平な公衆衛生研究を支援できる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。