[论文解读] Tracking and Quantifying Censorship on a Chinese Microblogging Site
本研究通过识别敏感用户并分析内容传播,追踪并量化了新浪微博这一中国社交媒体平台上的审查行为。利用递归社交爬取与自然语言处理技术,检测非正式、新造词汇的中文文本趋势,研究发现审查极为有效:敏感话题生命周期极短,且极少传播至核心用户群之外,实时与事后过滤机制有效抑制了病毒式传播。
We present measurements and analysis of censorship on Weibo, a popular microblogging site in China. Since we were limited in the rate at which we could download posts, we identified users likely to participate in sensitive topics and recursively followed their social contacts. We also leveraged new natural language processing techniques to pick out trending topics despite the use of neologisms, named entities, and informal language usage in Chinese social media. We found that Weibo dynamically adapts to the changing interests of its users through multiple layers of filtering. The filtering includes both retroactively searching posts by keyword or repost links to delete them, and rejecting posts as they are posted. The trend of sensitive topics is short-lived, suggesting that the censorship is effective in stopping the "viral" spread of sensitive issues. We also give evidence that sensitive topics in Weibo only scarcely propagate beyond a core of sensitive posters.
研究动机与目标
- 调查新浪微博这一主要中国社交媒体平台上的审查机制如何运作。
- 理解在国家强制内容过滤下,敏感话题传播的动态机制。
- 开发在非正式语言、新造词汇与快速话题更迭背景下,检测与追踪敏感话题的方法。
- 量化审查在限制敏感内容传播方面的有效性。
提出的方法
- 递归追踪可能发布敏感内容的用户社交网络,以在速率限制下扩展数据收集范围。
- 应用新颖的自然语言处理技术,识别非正式中文中的趋势话题,包括新造词与非标准表达。
- 监控基于关键词与链接的实时帖子拒收及事后删除机制。
- 追踪敏感话题在新浪微博用户网络中的生命周期与传播模式。
- 结合用户层级参与度指标与内容分析,评估审查的影响。
- 利用新浪微博API与受速率限制的网页爬取,收集帖子可见性与删除的纵向数据。
实验结果
研究问题
- RQ1在审查机制下,新浪微博上的敏感话题如何产生并传播?
- RQ2新浪微博采用何种机制抑制敏感内容——实时屏蔽还是事后删除?
- RQ3敏感话题在核心用户群之外的传播范围有多大?
- RQ4审查在多大程度上限制了新浪微博上敏感内容的病毒式传播?
- RQ5非正式语言与新造词汇在多大程度上影响了敏感话题的检测与追踪?
主要发现
- 新浪微博上的敏感话题生命周期极短,通常不足一天,表明其受到快速抑制。
- 审查通过多层机制运作:实时拒收帖子与事后删除已有内容。
- 仅一小部分紧密连接的核心用户群持续发布敏感话题内容,传播范围有限。
- 使用新造词与非正式语言并未显著阻碍所提出自然语言处理技术对趋势话题的检测。
- 系统能动态适应用户兴趣,表明其具备响应式与主动式的过滤策略。
- 尽管数据收集速率限制了研究范围,但通过针对性抽样,核心发现依然稳健。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。