Skip to main content
QUICK REVIEW

[论文解读] A personal model of trumpery: Deception detection in a real-world high-stakes setting

Sophie van der Zee, Ronald Poppe|arXiv (Cornell University)|Nov 5, 2018
Deception detection and forensic psychology参考文献 14被引用 3
一句话总结

本研究通过分析美国总统推文中的语言特征,并以《华盛顿邮报》的事实核查结果作为真实标签,开发了一种个性化语言模型,用于检测现实世界高风险沟通中的欺骗行为。该模型在预测样本外推文的事实正确性方面达到了73%的准确率,表明欺骗性语言模式在个体层面具有系统性和可识别性。

ABSTRACT

Language use reveals information about who we are and how we feel1-3. One of the pioneers in text analysis, Walter Weintraub, manually counted which types of words people used in medical interviews and showed that the frequency of first-person singular pronouns (i.e., I, me, my) was a reliable indicator of depression, with depressed people using I more often than people who are not depressed4. Several studies have demonstrated that language use also differs between truthful and deceptive statements5-7, but not all differences are consistent across people and contexts, making prediction difficult8. Here we show how well linguistic deception detection performs at the individual level by developing a model tailored to a single individual: the current US president. Using tweets fact-checked by an independent third party (Washington Post), we found substantial linguistic differences between factually correct and incorrect tweets and developed a quantitative model based on these differences. Next, we predicted whether out-of-sample tweets were either factually correct or incorrect and achieved a 73% overall accuracy. Our results demonstrate the power of linguistic analysis in real-world deception research when applied at the individual level and provide evidence that factually incorrect tweets are not random mistakes of the sender.

研究动机与目标

  • 探究社交媒体内容中的语言模式是否能可靠地在个体层面上预测事实准确性。
  • 为单一高知名度个体(特别是美国总统)开发一种个性化的欺骗检测模型。
  • 评估事实性错误的推文是否为随机错误,还是反映出系统性的语言模式。
  • 评估定量语言模型在预测现实世界高风险沟通中事实正确性的表现。

提出的方法

  • 从美国总统发布的1,000条推文中提取语言特征,这些推文由《华盛顿邮报》事实核查团队标记为事实正确或错误。
  • 使用监督式机器学习方法,基于这些标记过的推文训练个性化模型,以识别欺骗的语言标志。
  • 模型重点关注词汇、句法和语用特征,包括代词使用、情感倾向和结构复杂性。
  • 在未用于训练的样本外推文中测试该模型,以评估其泛化性能。
  • 使用标准分类指标衡量预测准确率,报告总体准确率为73%。

实验结果

研究问题

  • RQ1个性化语言模型能否以高准确率检测高风险社交媒体沟通中的欺骗行为?
  • RQ2同一个人的正确与错误事实推文之间是否存在一致的语言差异?
  • RQ3同一人发布的事实性错误推文是否表现出系统性语言模式,而非随机错误?
  • RQ4在单一人物上训练的模型在未见的同源推文上泛化能力如何?

主要发现

  • 个性化语言模型在预测推文事实正确性方面达到了73%的准确率。
  • 发现美国总统的正确与错误事实推文之间存在显著的语言差异。
  • 该模型表明,欺骗性推文并非随机错误,而是反映了语言使用中的一致性模式。
  • 结果支持在现实世界高风险场景中,通过语言分析实现个体层面欺骗检测的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。