[论文解读] Understanding the (In)Effectiveness of Content Moderation: A Case Study of Facebook in the Context of the U.S. Capitol Riot
本研究提出一种新方法,利用2020年1月6日国会山骚乱期间来自2,500家美国新闻来源的公开互动数据,推断Facebook内容删除的时间与影响。研究发现,即使删除速度很快,也仅能防止21%的预测互动量,因为大多数内容传播发生在30小时内——凸显了在危机中内容审核的系统性局限,并呼吁通过算法调整来减缓内容扩散。
Social media networks commonly employ content moderation as a tool to limit the spread of harmful content. However, the efficacy of this strategy in limiting the delivery of harmful content to users is not well understood. In this paper, we create a framework to quantify the efficacy of content moderation and use our metrics to analyze content removal on Facebook within the U.S. news ecosystem. In a data set of over 2M posts with 1.6B user engagements collected from 2,551 U.S. news sources before and during the Capitol Riot on January 6, 2021, we identify 10,811 removed posts. We find that the active engagement life cycle of Facebook posts is very short, with 90% of all engagement occurring within the first 30 hours after posting. Thus, even relatively quick intervention allowed significant accrual of engagement before removal, and prevented only 21% of the predicted engagement potential during a baseline period before the U.S. Capitol attack. Nearly a week after the attack, Facebook began removing older content, but these removals occurred so late in these posts' engagement life cycles that they disrupted less than 1% of predicted future engagement, highlighting the limited impact of this intervention. Content moderation likely has limits in its ability to prevent engagement, especially in a crisis, and we recommend that other approaches such as slowing down the rate of content diffusion be investigated.
研究动机与目标
- 评估Facebook内容审核在危机期间限制有害内容曝光的实际效果。
- 开发一种利用公开可获取的互动数据推断内容删除时间与影响的方法。
- 通过预测帖子病毒式传播的模型,量化删除前发生多少互动,以及实际防止了多少互动。
- 评估内容审核是否能在高互动事件(如国会山骚乱)中有效中断有害内容的传播。
- 建议提升透明度并采用替代策略(如减缓内容扩散),以增强平台安全性。
提出的方法
- 通过检测来自2,551家美国新闻来源的每日活跃Facebook帖子快照中帖子的消失,推断内容删除时间。
- 基于同一页面的历史互动模式,开发了预测帖子互动潜力的模型,对正常和病毒式传播帖子的误差率为3.2%至4.5%。
- 定义了两个关键指标:'累积互动量'(删除前的互动量)和'防止的互动量'(预测潜力减去实际互动量)。
- 使用基线期估计正常互动生命周期,发现90%的互动发生在发帖后30小时内。
- 将相同指标应用于危机期,比较1月12日Facebook政策变更前后删除时间与影响的差异。
- 利用第三方数据对新闻来源按事实准确性进行分类,以评估高风险内容是否被不成比例地删除。
实验结果
研究问题
- RQ1Facebook对美国新闻出版商和意见领袖的有害内容删除速度如何?删除前发生了多少互动?
- RQ2通过预测互动与实际互动的对比,内容审核在多大程度上防止了有害内容的曝光?
- RQ3在像1月6日国会山骚乱这样的高互动危机中,内容审核的有效性如何变化?
- RQ4删除时间相对于帖子互动生命周期的峰值有多晚?其实际影响是什么?
- RQ5哪些指标和透明度实践能够改善内容审核系统的设计与评估?
主要发现
- Facebook帖子的活跃互动生命周期极短,90%的互动发生在发帖后的前30小时内。
- 尽管中位删除时间为21小时,但在正常时期,内容审核仅防止了21.2%的预测互动潜力。
- 在国会山骚乱危机期间,初期删除速度缓慢且效果有限,仅防止了21%的预测互动,与基线水平相似。
- Facebook在袭击发生六天后开始删除旧内容,但这些延迟删除仅扰乱了不到1%的预测未来互动。
- 即使在政策变更和公告发布后,内容审核仍未能跟上危机期间内容扩散的速度。
- 结果表明,仅靠内容审核不足以防止有害内容的传播,应优先考虑替代策略(如减缓内容扩散)。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。