[论文解读] A Longitudinal Measurement Study of 4chan's Politically Incorrect Forum and its Effect on the Web.
本论文首次对4chan的/pol/论坛进行了纵向测量研究,分析了为期2.5个月内的超过800万条帖子,以描述全球用户分布、内容模式以及针对社交媒体的协调攻击。研究发现,/pol/推动了广泛传播的仇恨言论和YouTube链接分享,图像重复率极低,同时在通过语言混淆手段污染反网络欺凌工具方面成效有限。
The discussion board site 4chan has been a part of the dark underbelly of the Internet since its inception, but recent events have brought it to the forefront of the world's collective mind. In particular, /pol/, 4chan's Politically Incorrect board has become a central figure in the outlandish 2016 US election campaign, often linked to the alt-right movement and the rhetoric of hate and racism. Nonetheless, 4chan remains relatively unstudied by the research community. In this paper, we start addressing this gap by analyzing /pol/ along several axes, using a dataset of over 8M posts collected over two and a half months. First, we perform a general characterization, showing that /pol/ users are well distributed around the world and that 4chan's unique features encourage fresh discussions. Then, we analyze content posted on /pol/, finding YouTube links and hate speech to be predominant, that 95\% of images are posted no more than 5 times, and that there are notable differences in the English language used in different parts of the world. Last but not least, we provide quantitative evidence of /pol/'s collective attacks on other social media platforms by analyzing the comments in YouTube videos linked on /pol/. We also present a quantitative case study of /pol/'s attempt to poison anti-trolling tools by altering the language of hate on social media, finding it to be less successful than reported by the popular press. Overall, our analysis not only provides the first measurement study of /pol/, but also insight on online harassment and hate speech trends in online social media.
研究动机与目标
- 为解决关于/pol/——4chan上以政治不正确著称的论坛——缺乏实证研究的问题,该论坛在在线极端主义和2016年美国大选言论中日益突出。
- 分析/pol/用户的全球分布与语言多样性,评估4chan的结构特征如何促进新话题的产生。
- 调查仇恨言论、YouTube链接分享以及图像重复使用的普遍性,以理解内容动态。
- 量化/pol/通过评论垃圾信息攻击外部平台(尤其是YouTube)的协调攻击行为。
- 评估语言混淆技术在规避自动化反网络欺凌工具方面的有效性,挑战媒体广泛报道的成功说法。
提出的方法
- 通过网络爬虫和API访问,在为期2.5个月的时间内收集了超过800万条/posts/的纵向数据集。
- 对帖子进行地理定位和语言分析,以绘制用户分布图,并识别英语使用中的区域差异。
- 使用自然语言处理(NLP)技术检测仇恨言论,并对内容类型进行分类,包括YouTube链接的普遍性。
- 分析图像频率分布以评估重复使用模式,发现95%的图像被发布五次或更少次。
- 追踪从/pol/链接到的YouTube视频的评论,以衡量协调性踩低和骚扰活动。
- 通过比较语言操纵前后反网络欺凌工具的检测率,评估语言混淆策略的效果。
实验结果
研究问题
- RQ1 /pol/在地理上如何分布,4chan的结构特征在维持新话题讨论中起到什么作用?
- RQ2 /pol/中主导的内容类型是什么,特别是在仇恨言论、YouTube链接和图像分享方面?
- RQ3 /pol/在多大程度上通过评论垃圾信息对YouTube等外部平台实施协调攻击?
- RQ4 /pol/使用的语言混淆技术在规避自动化反网络欺凌工具方面有多有效?
- RQ5 /pol/中英语的区域语言差异如何反映全球用户的多样性?
主要发现
- /pol/的用户在全球范围内分布广泛,表明其用户基础具有跨国性质,而非局限于某一地区。
- 仇恨言论和YouTube链接是/pol/上最占主导地位的内容类型,主导了论坛的讨论。
- 95%的图像在/pol/上被发布五次或更少次,表明图像重复使用率低,内容更新频繁。
- 在英语使用上观察到显著的地理区域差异,反映出用户来源的多样性。
- 针对从/pol/链接到的YouTube评论的协调攻击可被定量测量,表明存在有组织的骚扰活动。
- 与媒体报道的结论相反,/pol/中使用的语言混淆技术在规避反网络欺凌工具方面效果有限,检测率依然很高。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。