Skip to main content
QUICK REVIEW

[論文レビュー] RAProp: Ranking Tweets by Exploiting the Tweet/User/Web Ecosystem and Inter-Tweet Agreement

Srijith Ravikumar, Kartik Talamadupula|arXiv (Cornell University)|Aug 11, 2013
Misinformation and Its Impacts参考文献 6被引用数 7
ひとこと要約

RAPropは、ユーザー、ツイート、ウェブページの特徴を用いたソース信頼性と、ツイート間相関に基づくコンテンツ合意を組み合わせる、新しいツイートランク手法である。信頼スコアは合意グラフ上で拡散され、関連性の向上に寄与する。TREC 2011 マイクロブログデータセットにおいて、現在の最良手法よりもトップ30の精度が53%高い。

ABSTRACT

The increasing popularity of Twitter renders improved trustworthiness and relevance assessment of tweets much more important for search. However, given the limitations on the size of tweets, it is hard to extract measures for ranking from the tweets' content alone. We present a novel ranking method, called RAProp, which combines two orthogonal measures of relevance and trustworthiness of a tweet. The first, called Feature Score, measures the trustworthiness of the source of the tweet. This is done by extracting features from a 3-layer twitter ecosystem, consisting of users, tweets and the pages referred to in the tweets. The second measure, called agreement analysis, estimates the trustworthiness of the content of the tweet, by analyzing how and whether the content is independently corroborated by other tweets. We view the candidate result set of tweets as the vertices of a graph, with the edges measuring the estimated agreement between each pair of tweets. The feature score is propagated over this agreement graph to compute the top-k tweets that have both trustworthy sources and independent corroboration. The evaluation of our method on 16 million tweets from the TREC 2011 Microblog Dataset shows that for top-30 precision we achieve 53% higher than current best performing method on the Dataset and over 300% over current Twitter Search. We also present a detailed internal empirical evaluation of RAProp in comparison to several alternative approaches proposed by us.

研究の動機と目的

  • ツイッター検索結果の関連性と信頼性の低さに起因する、リマインダー性とキーワードマッチングに依存する問題を解決すること。
  • TF-IDFのようなコンテンツ中心の手法の限界を克服し、短く情報量の少ないツイートに対して懲罰的措置を講じること。
  • ソースの信頼性とツイート内容の独立的裏付けを組み合わせることでランク付けを改善すること。
  • ツイート/ユーザー/ウェブエコシステムを活用した、スケーラブルでグラフベースの方法を開発し、堅牢な関連性評価を実現すること。
  • 本物のマイクロブログデータ上で、Twitter Search や USC/ISI といったベースラインを上回る精度と MAP メトリクスを達成すること。

提案手法

  • RAPropは、ユーザー、ツイート、ウェブページの3層構造を持つエコシステムモデルを構築し、ソース信頼性を反映する特徴スコア(FS)を計算する。
  • 頂点がツイート、エッジが共有コンテンツやウェブリファレンスに基づくツイート間のコンテンツ合意度を推定する、合意グラフを構築する。
  • 特徴スコアを合意グラフ上で拡散プロセスを用いて伝搬させ、裏付けのあるコンテンツを持つツイートの信頼スコアを向上させる。
  • FSの伝搬とツイート間合意を統合することで、信頼できるソースからのものであり、かつ独立的に裏付けられたツイートをランク付けする。
  • トップ-Kの精度と平均平均精度(MAP)を用いて結果を評価し、Twitter Search や USC/ISI といったベースラインと比較する。
  • 中間者モデルおよび非中間者モデルの両方で検証され、リtrieval設定にかかわらず堅牢性を示した。

実験結果

リサーチクエスチョン

  • RQ1ソース信頼性とツイート間コンテンツ合意を組み合わせることで、ツイートランクの精度が向上するか?
  • RQ2RAPropのグラフベース信頼スコア伝搬は、キーワードベースやリマインダーに基づくランク付けと比べて、関連性においてどのように差をつけるか?
  • RQ3ソース信頼性に加えて、コンテンツの独立的裏付けがどれほど信頼性を高めるか?
  • RQ4大規模マイクロブログデータ上での最先端手法(UC/ISI やネイティブな Twitter Search)と比較して、RAPropはどのように性能を発揮するか?
  • RQ5クエリ固有の背景知識に依存せずに、さまざまなトップ-Kの検索閾値においてもRAPropは高い精度を維持できるか?

主な発見

  • RAPropは、TREC 2011 マイクロブログデータセットにおいて、現在の最良手法よりもトップ30の精度が53%高い。
  • ネイティブな Twitter Search(リマインダー性とキーワードマッチングにのみ依存)と比較して、トップ30の精度が300%以上向上した。
  • 非中間者モデルにおいて、USC/ISI ベースラインよりもトップ30の精度が20%高い。
  • MAP において USC/ISI よりも4%高いスコアを記録し、より優れた全体的なランク品質を示した。
  • 合意グラフはコンテンツの裏付けを効果的に捉えており、信頼性が低い高リツイート数や高リマインダー性のツイートに依存するのを軽減した。
  • 合意グラフ上で特徴スコアを伝搬させることで、ソースが改ざんされたり誤解を招くものであっても、信頼できるツイートを効果的に同定できた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。