[论文解读] Modeling Islamist Extremist Communications on Social Media using Contextual Dimensions: Religion, Ideology, and Hate
本文提出了一种上下文感知的多维计算模型,通过宗教、意识形态和仇恨三个独立维度,分析Twitter上的伊斯兰极端主义内容,利用领域特定的知识库。通过将这些维度整合到随机森林和朴素贝叶斯分类器中,该方法相比基线模型将精确率提高了10.2%,显著减少了误标情况,提升了在反激进化行动中实际部署的可靠性。
Terror attacks have been linked in part to online extremist content. Although tens of thousands of Islamist extremism supporters consume such content, they are a small fraction relative to peaceful Muslims. The efforts to contain the ever-evolving extremism on social media platforms have remained inadequate and mostly ineffective. Divergent extremist and mainstream contexts challenge machine interpretation, with a particular threat to the precision of classification algorithms. Our context-aware computational approach to the analysis of extremist content on Twitter breaks down this persuasion process into building blocks that acknowledge inherent ambiguity and sparsity that likely challenge both manual and automated classification. We model this process using a combination of three contextual dimensions -- religion, ideology, and hate -- each elucidating a degree of radicalization and highlighting independent features to render them computationally accessible. We utilize domain-specific knowledge resources for each of these contextual dimensions such as Qur'an for religion, the books of extremist ideologues and preachers for political ideology and a social media hate speech corpus for hate. Our study makes three contributions to reliable analysis: (i) Development of a computational approach rooted in the contextual dimensions of religion, ideology, and hate that reflects strategies employed by online Islamist extremist groups, (ii) An in-depth analysis of relevant tweet datasets with respect to these dimensions to exclude likely mislabeled users, and (iii) A framework for understanding online radicalization as a process to assist counter-programming. Given the potentially significant social impact, we evaluate the performance of our algorithms to minimize mislabeling, where our approach outperforms a competitive baseline by 10.2% in precision.
研究动机与目标
- 为解决在社交媒体上准确识别伊斯兰极端主义用户的挑战,其中宗教语言常被扭曲以服务于极端主义意识形态。
- 克服现有分类方法的局限性,这些方法未能考虑极端主义传播中的语境模糊性和语义稀疏性。
- 开发一种稳健的多维框架,通过宗教、意识形态和仇恨维度建模在线激进化过程的渐进性和说服性特征。
- 通过整合领域特定的知识资源和上下文线索,减少对非极端主义用户误标的可能。
- 通过提供对极端主义叙事在不同激进化阶段演变的细致理解,支持有效的反宣传策略。
提出的方法
- 本研究构建了三个领域特定的知识资源:《古兰经》用于宗教,极端主义意识形态者的作品用于意识形态,社交媒体仇恨言论语料库用于仇恨。
- 通过三个上下文维度——宗教、意识形态和仇恨——对用户内容进行建模,每个维度均作为分类的独立特征空间。
- 该方法采用监督式机器学习,使用在标注了这些维度的推文数据集上训练的随机森林(RF)和朴素贝叶斯(NB)分类器。
- 特征工程包括从领域特定语料库中提取的上下文嵌入,以捕捉宗教、意识形态和仇恨表达中的语义细微差别。
- 通过精确率、召回率、F1分数和AUC评估模型性能,并将多维模型与单维基线模型进行比较。
- 采用排除策略,基于三个上下文维度之间不一致的对齐情况,移除可能被误标的用户。
实验结果
研究问题
- RQ1在社交媒体上的伊斯兰极端主义传播中,宗教、意识形态和仇恨如何作为既独立又相互关联的维度发挥作用?
- RQ2与单维方法相比,整合多个上下文维度在多大程度上提升了分类极端主义用户的精确率和可靠性?
- RQ3宗教、意识形态和仇恨在在线激进化不同阶段中的相对贡献如何变化?
- RQ4上下文感知模型能否减少对使用宗教语言但无极端主义意图的非极端主义用户的误分类?
- RQ5在Twitter上检测极端主义内容方面,多维建模相比传统基线方法的性能提升有多大?
主要发现
- 三维度模型相比基线模型将精确率提高了10.2%,显著降低了将非极端主义用户误标的潜在风险。
- 结合宗教和仇恨的双维度模型(RH)达到最高的AUC值0.90,优于三维度模型及其他双维度模型。
- 三维度随机森林模型在假正例率(FPR)为0.65时达到0.93的真正例率(TPR),表明在高精确率约束下表现强劲。
- 包含全部三个维度的模型实现了最佳整体性能,相比基线模型,召回率提升8.5%,F1分数提升10.7%。
- 三维度朴素贝叶斯模型的ROC曲线在FPR为0.83时达到1.0的TPR,而双维度模型在FPR为0.97时达到1.0的TPR,表明精确率提升了9.8%。
- 宗教语言和仇恨语言被发现比仅依赖意识形态更具区分力,表明它们在极端主义正当化叙事中常共同出现。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。