Skip to main content
QUICK REVIEW

[论文解读] An Analysis of Chinese Search Engine Filtering

Tao Zhu, Chris Bronk|arXiv (Cornell University)|Jul 19, 2011
Internet Traffic Analysis and Secure E-voting参考文献 25被引用 4
一句话总结

本文通过在百度、谷歌(终止前)、雅虎和必应上测试查询,分析了中国搜索引擎的过滤机制,揭示了对色情、政治及活动人士相关词汇的激烈关键词过滤。研究识别出动态演变的过滤政策——通常涉及黑名单和政府控制的白名单——展示了搜索引擎如何在中国执行国家规定的网络内容限制。

ABSTRACT

The imposition of government mandates upon Internet search engine operation is a growing area of interest for both computer science and public policy. Users of these search engines often observe evidence of censorship, but the government policies that impose this censorship are not generally public. To better understand these policies, we conducted a set of experiments on major search engines employed by Internet users in China, issuing queries against a variety of different words: some neutral, some with names of important people, some political, and some pornographic. We conducted these queries, in Chinese, against Baidu, Google (including google.cn, before it was terminated), Yahoo!, and Bing. We found remarkably aggressive filtering of pornographic terms, in some cases causing non-pornographic terms which use common characters to also be filtered. We also found that names of prominent activists and organizers as well as top political and military leaders, were also filtered in whole or in part. In some cases, we found search terms which we believe to be "blacklisted". In these cases, the only results that appeared, for any of them, came from a short "whitelist" of sites owned or controlled directly by the Chinese government. By repeating observations over a long observation period, we also found that the keyword blocking policies of the Great Firewall of China vary over time. While our results don't offer any fundamental insight into how to defeat or work around Chinese internet censorship, they are still helpful to understand the structure of how censorship duties are shared between the Great Firewall and Chinese search engines.

研究动机与目标

  • 调查在中国国家指令下,搜索引擎内容过滤的范围与机制。
  • 识别哪些类型的搜索关键词——尤其是政治、活动人士或色情类——被系统性屏蔽。
  • 考察搜索引擎在执行“防火墙”审查政策中的角色。
  • 记录过滤行为随时间的变化情况。
  • 厘清“防火墙”与搜索引擎运营商在内容压制中的职责分工。

提出的方法

  • 使用多样化关键词集合(中性词、政治类词汇、知名人物姓名、色情词汇)在中文环境下进行受控查询。
  • 测试了四大主流搜索引擎:百度、谷歌(撤出前)、雅虎和必应。
  • 在较长观察期内重复查询,以检测过滤行为的时间变化。
  • 分析返回结果,识别抑制模式,如结果完全缺失或仅显示政府批准的网站。
  • 通过观察特定关键词在多个查询中的一致性过滤,识别潜在黑名单。
  • 区分广义字符过滤(如非色情词汇中常见的字符)与针对性关键词屏蔽。

实验结果

研究问题

  • RQ1哪些类别的搜索关键词最常被中国搜索引擎过滤?
  • RQ2搜索引擎在多大程度上通过黑名单和白名单实施审查?
  • RQ3过滤政策如何随时间变化?其变化背后可能受哪些因素驱动?
  • RQ4非色情词汇是否因与被屏蔽词汇共享字符而被过滤?
  • RQ5搜索引擎在执行国家级网络审查中的角色是什么,与“防火墙”本身相比如何?

主要发现

  • 色情词汇受到强烈过滤,部分使用常见字符的非色情词汇也一并被屏蔽。
  • 知名政治人物、军事领导人及活动人士的姓名被系统性地过滤或压制。
  • 对于某些黑名单词汇,结果中仅显示少量政府所有或控制的网站。
  • 过滤政策表现出时间上的变化,表明其执行策略是动态且持续演化的。
  • 本研究证实,中国搜索引擎积极参与国家主导的审查,常作为“防火墙”与终端用户之间的中介。
  • 结果表明,过滤并非纯粹技术行为,而是反映明确的政策选择,部分词汇即使未被明确定义为非法,仍被屏蔽。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。