[论文解读] How Curiosity can be modeled for a Clickbait Detector
本文提出了一种新颖的计算模型,利用心理学和认知理论,结合手工设计的与新颖性、意外性和信息缺口相关的特征,通过无监督与有监督机器学习相结合的方法,量化点击诱饵标题中的人类好奇心。该模型将好奇心建模为注意力的驱动力,在分类点击诱饵标题方面实现了97.17%的准确率,是首次在数字内容检测中实现好奇心的可操作化。
The impact of continually evolving digital technologies and the proliferation of communications and content has now been widely acknowledged to be central to understanding our world. What is less acknowledged is that this is based on the successful arousing of curiosity both at the collective and individual levels. Advertisers, communication professionals and news editors are in constant competition to capture attention of the digital population perennially shifty and distracted. This paper, tries to understand how curiosity works in the digital world by attempting the first ever work done on quantifying human curiosity, basing itself on various theories drawn from humanities and social sciences. Curious communication pushes people to spot, read and click the message from their social feed or any other form of online presentation. Our approach focuses on measuring the strength of the stimulus to generate reader curiosity by using unsupervised and supervised machine learning algorithms, but is also informed by philosophical, psychological, neural and cognitive studies on this topic. Manually annotated news headlines - clickbaits - have been selected for the study, which are known to have drawn huge reader response. A binary classifier was developed based on human curiosity (unlike the work done so far using words and other linguistic features). Our classifier shows an accuracy of 97% . This work is part of the research in computational humanities on digital politics quantifying the emotions of curiosity and outrage on digital media.
研究动机与目标
- 为解决数字内容中缺乏人类好奇心的计算模型的问题,将方法建立在心理学和认知理论基础上。
- 开发一种二元分类器,基于标题激发好奇心的能力来区分点击诱饵标题,而非仅依赖语言或词汇特征。
- 验证好奇心——尤其是认识性和感知性形式——可以被量化并用于预测在线内容的用户参与度。
- 探讨标题中的风格、语义和结构特征如何通过信息缺口、新颖性和意外性触发好奇心。
- 通过实现对数字媒体中情绪敏感的分析,特别是在线话语中好奇心与愤怒情绪的分析,推动计算人文学的发展。
提出的方法
- 本研究结合无监督与有监督机器学习,基于心理学中关于好奇心的理论(包括新颖性、意外性和信息缺口)所提取的特征,对点击诱饵标题进行分类。
- 特征基于认知与神经科学研究手工设计:新颖性建模为与常规主题的语义偏离,意外性建模为非常规词序或句法破坏,信息缺口建模为暗示知识不完整的修辞线索。
- 应用主题模型(LDA)识别标题的主题新颖性,使用一致性分数选择最优主题数(N=68),证实点击诱饵的主题集中于身份、自我和流行文化。
- 模型采用SVM与逻辑回归分类器,在1:4的训练-测试划分上进行训练,性能通过准确率、F1-score和均方误差(MSE)进行评估。
- 该方法整合了I&D模型(Litman, 2005)的研究成果,其中I-interest(认识性)与D-deprivation(感知性)好奇心被视为不同但互补的参与驱动力。
- 通过期望违背与知识缺口评估刺激物引发认知唤醒的能力,结果与多巴胺对信息的神经相关性保持一致。
实验结果
研究问题
- RQ1如何在数字标题中计算建模人类好奇心,特别是认识性和感知性形式?
- RQ2新颖性、意外性和信息缺口等特征在多大程度上比传统语言特征更有效地预测点击诱饵的有效性?
- RQ3基于好奇心驱动特征训练的机器学习分类器是否能优于现有点击诱饵检测方法?
- RQ4点击诱饵标题的主题分布与普通新闻标题相比,在语义新颖性和与自我身份的相关性方面有何差异?
- RQ5语义新颖性、句法意外性和修辞信息缺口在驱动读者好奇心方面的相对贡献是什么?
主要发现
- 使用所有手工设计特征(新颖性、意外性、信息缺口)的模型实现了97.17%的分类准确率,显著优于仅基于新颖性或意外性的模型。
- 当仅使用新颖性特征时,模型准确率最高(99.17%),表明与常规主题的语义偏离是点击诱饵潜力最强的预测因子。
- 主题模型显示,点击诱饵标题显著更可能涉及与自我、身份、名人及流行文化(如占星术、美容、健康)相关的话题,而主流新闻则多涉及政治或战争等主题。
- 点击诱饵标题通过列表式结构(使用数字)、反问句和不完整陈述等风格特征制造强烈的信息缺口,促使读者寻求认知闭合。
- 基于意外性的特征准确率较低(93.69%)但F1-score较高(0.9674),表明句法或词汇上的不可预测性是可靠但不如语义新颖性占主导地位的信号。
- 结果支持理论框架:感知性好奇心(由厌恶驱动)与认识性好奇心(由愉悦驱动)是不同但互补的机制,其中认识性好奇心与高性能点击诱饵特征的匹配度更强。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。