[論文レビュー] A personal model of trumpery: Deception detection in a real-world high-stakes setting
本研究では、ウォ싱トン・ポストのfact-checkを基準事実として用い、米国大統領のツイートを分析することで、現実世界の高リスクなコミュニケーションにおける不正行為を検出するためのパーソナライズドな言語モデルを開発した。アウトオブサンプルのツイートについて、事実の正しさを予測する精度は73%に達し、不正な言語パターンが個人レベルで体系的かつ識別可能であることを示している。
Language use reveals information about who we are and how we feel1-3. One of the pioneers in text analysis, Walter Weintraub, manually counted which types of words people used in medical interviews and showed that the frequency of first-person singular pronouns (i.e., I, me, my) was a reliable indicator of depression, with depressed people using I more often than people who are not depressed4. Several studies have demonstrated that language use also differs between truthful and deceptive statements5-7, but not all differences are consistent across people and contexts, making prediction difficult8. Here we show how well linguistic deception detection performs at the individual level by developing a model tailored to a single individual: the current US president. Using tweets fact-checked by an independent third party (Washington Post), we found substantial linguistic differences between factually correct and incorrect tweets and developed a quantitative model based on these differences. Next, we predicted whether out-of-sample tweets were either factually correct or incorrect and achieved a 73% overall accuracy. Our results demonstrate the power of linguistic analysis in real-world deception research when applied at the individual level and provide evidence that factually incorrect tweets are not random mistakes of the sender.
研究の動機と目的
- SNS投稿における言語的パターンが、個人レベルで事実の正しさを信頼性を持って予測できるかどうかを調査すること。
- 特に米国大統領を対象として、1人の著名な個人に特化した不正検出のパーソナライズドモデルを構築すること。
- 事実誤認のツイートがランダムな誤りであるのか、それとも体系的な言語的パターンを示すのかを評価すること。
- 定量的言語的モデルが、現実世界の高リスクなコミュニケーションにおける事実の正しさを予測する性能を評価すること。
提案手法
- ウォ싱トン・ポストのfact-checkチームが事実的正誤をラベル付けした1,000件の米国大統領のツイートから、言語的特徴を抽出した。
- 教師あり機械学習を用いて、これらのラベル付きツイート上で個人用モデルを訓練し、不正の兆候となる言語的マーカーを同定した。
- モデルは、語彙的・構文的・意味的特徴に焦点を当て、代名詞の使用、感情、構造の複雑さを含めた。
- 一般化性能を評価するために、訓練に使用しなかったアウトオブサンプルのツイートでモデルをテストした。
- 分類の標準的指標を用いて精度を測定し、73%の全体精度が報告された。
実験結果
リサーチクエスチョン
- RQ1パーソナライズドな言語的モデルは、高リスクなSNSコミュニケーションにおける不正を高い精度で検出できるか?
- RQ21人の個人からの事実的正誤ツイート間に一貫した言語的差異が存在するか?
- RQ3同じ個人が発信する事実誤認ツイートは、ランダムな誤りではなく体系的な言語的パターンを示すのか?
- RQ41人の個人で訓練されたモデルは、同じ出処の未観測ツイートにどの程度一般化できるか?
主な発見
- パーソナライズドな言語的モデルは、ツイートが事実的正しくないかどうかを73%の精度で予測した。
- 米国大統領の事実的正誤ツイート間には顕著な言語的差異が確認された。
- モデルは、不正なツイートがランダムなミスではなく、一貫した言語使用のパターンを示していることを示した。
- 結果は、現実世界の高リスクな状況において、言語的分析を用いた個人レベルの不正検出の可能性を支持する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。