[论文解读] Multimodal Social Media Analysis for Gang Violence Prevention
本文提出一种多模态方法,通过结合文本和图像数据,检测芝加哥帮派关联青少年社交媒体帖子中的心理社会编码(攻击性、丧失感、物质滥用)。通过融合文本与图像模态,该方法在平均平均精度上相比单模态模型实现了18%的相对提升,证明了多模态分析在预防帮派暴力早期干预中的价值。
Gang violence is a severe issue in major cities across the U.S. and recent studies [Patton et al. 2017] have found evidence of social media communications that can be linked to such violence in communities with high rates of exposure to gang activity. In this paper we partnered computer scientists with social work researchers, who have domain expertise in gang violence, to analyze how public tweets with images posted by youth who mention gang associations on Twitter can be leveraged to automatically detect psychosocial factors and conditions that could potentially assist social workers and violence outreach workers in prevention and early intervention programs. To this end, we developed a rigorous methodology for collecting and annotating tweets. We gathered 1,851 tweets and accompanying annotations related to visual concepts and the psychosocial codes: aggression, loss, and substance use. These codes are relevant to social work interventions, as they represent possible pathways to violence on social media. We compare various methods for classifying tweets into these three classes, using only the text of the tweet, only the image of the tweet, or both modalities as input to the classifier. In particular, we analyze the usefulness of mid-level visual concepts and the role of different modalities for this tweet classification task. Our experiments show that individually, text information dominates classification performance of the loss class, while image information dominates the aggression and substance use classes. Our multimodal approach provides a very promising improvement (18% relative in mean average precision) over the best single modality approach. Finally, we also illustrate the complexity of understanding social media data and elaborate on open challenges.
研究动机与目标
- 开发一个严谨的框架,用于收集并标注来自芝加哥帮派关联青少年的公开推文(含图片)。
- 识别并标注与社会工作干预及暴力路径相关的心理社会编码(攻击性、丧失感、物质滥用)。
- 评估仅使用文本、仅使用图像以及多模态方法在分类这些心理社会编码方面的有效性。
- 分析中层视觉概念及各模态特异性贡献在提升分类性能中的作用。
- 通过移除标识符、加密数据以及引入社区专家参与验证,确保数据处理的伦理合规性。
提出的方法
- 从芝加哥帮派关联青少年处收集了1,851条含文本内容及关联图片的公开推文。
- 由社会工作研究人员与领域专家组成的团队,对每条推文进行心理社会编码(攻击性、丧失感、物质滥用)及视觉概念标注。
- 使用深度学习模型提取视觉与文本特征,训练并比较仅使用文本、仅使用图像或双模态的分类器。
- 将局部视觉概念(如手势、武器、毒品用具)与全局图像特征相结合,以提升对攻击性与物质滥用的检测能力。
- 采用多模态融合策略整合文本与图像预测结果,提升整体分类性能。
- 通过领域专家的迭代验证解决标注分歧,尤其针对文化背景与象征性手势的理解。
实验结果
研究问题
- RQ1心理社会编码(攻击性、丧失感、物质滥用)在帮派关联青少年的多模态社交媒体内容中如何表现?
- RQ2在检测三种心理社会编码时,文本与图像模态的相对贡献分别是什么?
- RQ3中层视觉概念在多大程度上提升了对推文中攻击性与物质滥用的检测能力?
- RQ4多模态融合方法与单模态方法相比,在分类心理社会编码方面表现如何?
- RQ5在高暴力社区中,标注与解读文化内涵丰富的社交媒体内容面临哪些主要挑战?
主要发现
- 仅使用文本的分类在‘丧失感’编码上表现优于仅使用图像的分类,表明文本线索在检测情感困扰方面更具可靠性。
- 仅使用图像的分类在‘攻击性’与‘物质滥用’编码上表现优于仅使用文本的分类,表明视觉线索在检测暴力或高风险行为方面更具信息量。
- 多模态融合方法在平均平均精度上相比最佳单模态方法实现了18%的相对提升,表明两种模态之间存在显著协同效应。
- 标注者与领域专家之间的分歧主要源于文化知识的缺乏,特别是对帮派归属、本地地标及象征性手势的理解不足。
- 领域专家对象征性手势(如表示不敬的特定手势)的理解显著改变了攻击性评估结果,凸显了语境专业知识的重要性。
- 尽管自动分类性能显著提升,但文化细微差别仍构成主要限制,强调了人类专家在标注与验证中的不可替代作用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。