[论文解读] Documenting the English Colossal Clean Crawled Corpus.
本文首次全面记录了大规模、经过筛选的文本语料库——巨量清洁爬取语料库(C4)的详细信息。研究分析了数据来源、过滤过程的影响(尤其是对非洲裔美国人英语,AAE 的影响),并识别出语料库中包含的基准数据集示例,同时发布了交互式网络界面,供社区探索使用。
As language models are trained on ever more text, researchers are turning to some of the largest corpora available. Unlike most other types of datasets in NLP, large unlabeled text corpora are often presented with minimal documentation, and best practices for documenting them have not been established. In this work we provide the first documentation for the Colossal Clean Crawled Corpus (C4; Raffel et al., 2020), a dataset created by applying a set of filters to a single snapshot of Common Crawl. We begin with a high-level summary of the data, including distributions of where the text came from and when it was written. We then give more detailed analysis on salient parts of this data, including the most frequent sources of text (e.g., this http URL, which contains a significant percentage of machine translated and/or OCR'd text), the effect that the filters had on the data (they disproportionately remove text in AAE), and evidence that some other benchmark NLP dataset examples are contained in the text. We release a web interface to an interactive, indexed copy of this dataset, encouraging the community to continuously explore and report additional findings.
研究动机与目标
- 建立自然语言处理(NLP)领域中大型无标注文本语料库的文档规范,以解决当前缺乏标准化文档的问题。
- 对巨量清洁爬取语料库(C4)进行详细分析,包括数据来源、时间分布及来源特征。
- 研究 C4 中使用的过滤流程对语言多样性的影响,特别是对非洲裔美国人英语(AAE)文本的不成比例影响。
- 识别并验证已知基准 NLP 数据集示例在 C4 语料库中的存在。
- 发布一个交互式、索引化的网络界面,以支持对语料库的持续社区探索与发现。
提出的方法
- 作者分析 Common Crawl 数据集的一个快照,以追踪 C4 中文本的来源及其时间分布。
- 识别并分类最常见的文本来源,包括包含高比例机器翻译或 OCR 处理内容的 URL。
- 评估过滤流程对语言多样性的影响,特别是其对非洲裔美国人英语(AAE)文本的影响。
- 系统性地搜索以检测 C4 语料库中是否存在已知的基准 NLP 示例(如 GLUE、SuperGLUE)。
- 构建并发布一个网络界面,对 C4 语料库进行索引,以支持交互式、实时探索和社区驱动的发现。
实验结果
研究问题
- RQ1巨量清洁爬取语料库中的文本主要来自哪些来源,其时间分布如何?
- RQ2C4 中使用的过滤流程如何影响特定语言变体(如非洲裔美国人英语,AAE)的代表性?
- RQ3现有基准 NLP 数据集在 C4 语料库中有多大程度上存在?
- RQ4在 C4 中最常见的 URL 中,哪些类型的文本(如机器翻译、OCR 处理)占主导地位?
- RQ5交互式、索引化的界面在提升社区参与度和促进对大型文本语料库的持续分析方面有何作用?
主要发现
- C4 中最常见的来源是一个单一 URL,其中包含高比例的机器翻译和/或 OCR 处理的文本。
- 过滤流程对非洲裔美国人英语(AAE)文本的去除存在不成比例的情况,导致其在最终语料库中的代表性降低。
- 研究发现,已知的基准 NLP 数据集示例(如 GLUE 和 SuperGLUE 中的示例)存在于 C4 语料库中。
- 该数据集包含多样化的文本来源,其中相当大一部分源自低质量或自动化内容的网页爬取。
- 发布交互式、索引化的网络界面,支持实时探索,并促进社区对语料库特征的发现。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。