[论文解读] Feels Bad Man: Dissecting Automated Hateful Meme Detection Through the Lens of Facebook's Challenge
本研究利用来自 4chan 的 /pol/ 板块和 Facebook 的 Hateful Memes Challenge 的数据集,评估了当前最先进的多模态模型在仇恨表情包检测中的表现。研究发现,仅依靠视觉特征的模型表现优于纯文本分析方法,且在合成的 Facebook 数据上训练的模型在真实世界中的边缘平台病毒式传播表情包上泛化能力极差,凸显了当前检测系统的关键局限性。
Internet memes have become a dominant method of communication; at the same time, however, they are also increasingly being used to advocate extremism and foster derogatory beliefs. Nonetheless, we do not have a firm understanding as to which perceptual aspects of memes cause this phenomenon. In this work, we assess the efficacy of current state-of-the-art multimodal machine learning models toward hateful meme detection, and in particular with respect to their generalizability across platforms. We use two benchmark datasets comprising 12,140 and 10,567 images from 4chan's "Politically Incorrect" board (/pol/) and Facebook's Hateful Memes Challenge dataset to train the competition's top-ranking machine learning models for the discovery of the most prominent features that distinguish viral hateful memes from benign ones. We conduct three experiments to determine the importance of multimodality on classification performance, the influential capacity of fringe Web communities on mainstream social platforms and vice versa, and the models' learning transferability on 4chan memes. Our experiments show that memes' image characteristics provide a greater wealth of information than its textual content. We also find that current systems developed for online detection of hate speech in memes necessitate further concentration on its visual elements to improve their interpretation of underlying cultural connotations, implying that multimodal models fail to adequately grasp the intricacies of hate speech in memes and generalize across social media platforms.
研究动机与目标
- 评估多模态机器学习模型在不同社交媒体平台中检测仇恨表情包的有效性。
- 探究在 Facebook 提供的合成数据集上训练的模型在真实世界中的边缘社区(如 4chan 的 /pol/)病毒式传播表情包上的泛化能力。
- 识别区分病毒式传播仇恨表情包与良性表情包的关键视觉与文本特征。
- 评估多模态性(图像与文本)对分类性能及模型泛化能力的影响。
- 揭示现有数据集与模型在多模态表情包仇恨言论检测中的局限性。
提出的方法
- 在来自 4chan 的 /pol/ 板块的 12,140 张表情包数据集上训练 VisualBERT,以评估文本内容对分类的贡献。
- 在 4chan 表情包上评估经过 Facebook Hateful Memes Challenge 数据集预训练的 UNITER 模型,以测试跨平台泛化能力。
- 仅在 4chan 表情包上训练三种模型——UNITER、OSCAR 和一个集成模型,以评估顶级挑战解决方案的可迁移性。
- 对表现最佳的模型进行特征重要性分析,以识别在仇恨表情包分类中最具影响力的视觉属性。
- 使用 Kiela 等人和 Zannettou 等人的基准数据集,以确保实验间的一致性与可比性。
- 应用单模态(仅图像)与多模态(图像 + 文本)表示,以隔离视觉与文本特征的影响。
实验结果
研究问题
- RQ1多模态性在图像表情包的仇恨表情包检测中有多大的影响?
- RQ2在 Facebook Hateful Memes Challenge 数据集上训练的模型在其他平台(如 4chan)的表情包上有多强的可移植性?
- RQ3区分病毒式传播仇恨表情包与良性表情包的关键视觉与文本特征是什么?
- RQ4当前多模态模型在跨平台场景下的泛化能力如何,尤其是在使用合成数据训练的情况下?
- RQ5哪些视觉属性与表情包的病毒式传播性及仇恨感知度最强相关?
主要发现
- 表情包的视觉特征提供的区分性信息显著多于文本内容,即使在单模态(仅图像)设置下也能实现 80% 的分类准确率。
- 在 Facebook 合成 Hateful Memes Challenge 数据集上训练的模型在真实世界 4chan 表情包上的表现较差,表明其泛化能力弱,且基准数据集代表性不足。
- 表现最佳的集成模型在分类 4chan 的病毒式传播仇恨表情包时达到了 84% 的准确率,证明了多模态融合在真实世界检测中的有效性。
- 四个关键视觉属性——主题内容、面部表情、手势和比例——在表情包的病毒式传播性与仇恨感知度方面始终具有显著关联。
- 仅依靠文本内容无法充分解释分类性能,因为模型在不依赖文本特征的情况下也能实现高准确率。
- 研究发现模型预测存在偏差,尤其对“Jew”等词汇在非仇恨语境下也容易被标记为仇恨内容,凸显了仇恨检测中文化与语境细微差别的挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。