Skip to main content
QUICK REVIEW

[論文レビュー] Mining Anonymity: Identifying Sensitive Accounts on Twitter

Sai Teja Peddinti, Keith W. Ross|arXiv (Cornell University)|Feb 1, 2017
Spam and Phishing Detection参考文献 23被引用数 6
ひとこと要約

本稿では、キーワードベースの検出に依存せずに、フォロワーの匿名性パターンを分析することで、スケーラブルで客観的な方法で、Twitterのセンシティブなアカウントを特定することを提案する。機械学習分類器を用いて匿名フォロワーと識別可能なフォロワーを区別するように訓練することで、匿名フォロワーの割合が高いアカウント(=センシティブ性を示す)を同定する。LDAトピックモデリングと人間評価を用いた検証により、抑圧されたトピック(例:うつ病、LGBTQ関連)に関するアカウントを効果的に特定した。

ABSTRACT

We explore the feasibility of automatically finding accounts that publish sensitive content on Twitter. One natural approach to this problem is to first create a list of sensitive keywords, and then identify Twitter accounts that use these words in their tweets. But such an approach may overlook sensitive accounts that are not covered by the subjective choice of keywords. In this paper, we instead explore finding sensitive accounts by examining the percentage of anonymous and identifiable followers the accounts have. This approach is motivated by an earlier study showing that sensitive accounts typically have a large percentage of anonymous followers and a small percentage of identifiable followers. To this end, we first considered the problem of automatically determining if a Twitter account is anonymous or identifiable. We find that simple techniques, such as checking for name-list membership, perform poorly. We designed a machine learning classifier that classifies accounts as anonymous or identifiable. We then classified an account as sensitive based on the percentages of anonymous and identifiable followers the account has. We applied our approach to approximately 100,000 accounts with 404 million active followers. The approach uncovered accounts that were sensitive for a diverse number of reasons. These accounts span across varied themes, including those that are not commonly proposed as sensitive or those that relate to socially stigmatized topics. To validate our approach, we applied Latent Dirichlet Allocation (LDA) topic analysis to the tweets in the detected sensitive and non-sensitive accounts. LDA showed that the sensitive and non-sensitive accounts obtained from the methodology are tweeting about distinctly different topics. Our results show that it is indeed possible to objectively identify sensitive accounts at the scale of Twitter.

研究の動機と目的

  • キーワードベースの検出における主観性を回避する、客観的でスケーラブルなセンシティブなTwitterアカウントの特定手法を開発すること。
  • フォロワーの匿名性パターン(特に、匿名フォロワーと識別可能フォロワーの割合)が、コンテンツのセンシティブ性を信頼できる代理指標として機能するかどうかを調査すること。
  • トピックモデリング(LDA)と人間評価を用いて手法を検証し、同定されたアカウントが実際にセンシティブなテーマについて議論していることを保証すること。
  • 本手法が、事前に定義されたキーワードリストでは捉えきれない多様な社会的ステグマを伴うトピックにも一般化可能であることを示すこと。

提案手法

  • プロフィール特徴(ユーザーネーム、バイオ、メタデータなど)に基づいて、Twitterアカウントが匿名か識別可能かを自動的に特定する機械学習分類器を訓練する。
  • 分類器を用いて、各ターゲットアカウントの匿名フォロワーと識別可能フォロワーの割合を算出する。
  • 匿名フォロワーの割合が高く、識別可能フォロワーの割合が低い場合に、アカウントをセンシティブと分類する。
  • 約10万件のアカウントと4億400万件のアクティブフォロワーを含む大規模なTwitterクロールに本手法を適用する。
  • 2,000件のセンシティブおよび非センシティブアカウントのサブセットに対して、Latent Dirichlet Allocation (LDA) トピックモデリングを実施して結果を検証する。
  • 200件のアカウントについて人間評価を実施し、自動分類と人間の判断との整合性を評価する。

実験結果

リサーチクエスチョン

  • RQ1フォロワーの匿名性パターンは、Twitter上でのセンシティブなコンテンツを信頼できる客観的指標として機能するか?
  • RQ2本手法は、従来のキーワードベースのアプローチではカバーできないセンシティブなトピックにも一般化可能か?
  • RQ3匿名性パターンに基づく自動分類は、コンテンツのセンシティブ性についての人間の判断とどの程度整合性を示すか?
  • RQ4本手法で同定されたセンシティブなアカウントは、非センシティブなアカウントと比較して、明確に特徴的なトピック的テーマを扱っているか?

主な発見

  • LDAトピック分析の結果、本手法で同定されたセンシティブなアカウントは、非センシティブなアカウントと明確に異なるトピックを扱っていることが判明し、トピックの差異化が確認された。
  • 人間評価者との間に強い一致が認められ、大多数のセンシティブなアカウントが、うつ病、LGBTQ関連、ポルノグラフィーなどのトピックに基づいて正しく分類された。
  • 本手法は、キーワードベースのシステムがしばしば見過ごす、抑圧的または社会的に敏感なトピックについて議論するアカウントを効果的に特定できた。
  • 本手法は、言語固有のキーワードに依存せずに、全Twitterエコシステムにスケーラブルに適用可能であり、約10万件のアカウントと4億400万件のフォロワーを処理できた。
  • 匿名フォロワーの割合が高いアカウントは、一貫してセンシティブなコンテンツと関連づけられており、本手法の核心的仮定が裏付けられた。
  • 本研究では、フォロワーの匿名性パターンが、多様で微細なトピックに対しても、堅牢でスケーラブルなコンテンツセンシティブ性の代理指標として機能可能であることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。