[Paper Review] Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models
This paper introduces CoWeSe, the largest publicly available Spanish biomedical corpus to date, comprising 4.5GB of cleaned plain text (745.7 million tokens) from 2.766 websites across 3,338 curated domains. The corpus was created via large-scale web crawling (2020) with domain-specific seeding and a custom preprocessing pipeline, enabling high-quality, diverse biomedical text for training Spanish language models and advancing NLP in health and biomedicine.
We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result of a massive crawler on 3000 Spanish domains executed in 2020. The corpus is openly available and already preprocessed. CoWeSe is an important resource for biomedical and health NLP in Spanish and has already been employed to train domain-specific language models and to produce word embbedings. We released the CoWeSe corpus under a Creative Commons Attribution 4.0 International license, both in Zenodo (\url{https://zenodo.org/record/4561971\#.YTI5SnVKiEA}).
Motivation & Objective
- To address the scarcity of large-scale, high-quality Spanish biomedical text for NLP by creating a comprehensive, diverse corpus.
- To overcome limitations in existing resources by collecting data from non-academic, real-world health websites beyond scientific literature.
- To support the development of domain-specific language models and word embeddings for Spanish biomedical and clinical NLP applications.
- To provide a publicly accessible, preprocessed corpus that enables reproducible research and model training in Spanish biomedicine.
Proposed method
- Crawled 3,338 curated websites across 10 health-related categories, including hospitals, patient associations, and public health bodies, using a depth-5 crawl strategy.
- Selected domains were prioritized based on relevance to health and biomedical content, including sites designated by Spain’s ISCIII as health information resources.
- Extracted only paragraph and header HTML tags to preserve meaningful biomedical text while excluding non-relevant formats.
- Applied a customized high-performance computing pipeline (50-node cluster) for cleaning, including sentence splitting, language identification, and deduplication.
- Increased language identification thresholds to better detect biomedical Spanish, which often has lower confidence scores in standard models.
- Preserved original document boundaries and applied content deduplication to ensure data quality and prevent redundancy.
Experimental results
Research questions
- RQ1Can a large-scale, web-crawled corpus provide sufficient coverage and diversity for training effective Spanish biomedical language models?
- RQ2How does a corpus derived from non-academic, real-world health websites compare to traditional scientific literature corpora in terms of linguistic and domain diversity?
- RQ3To what extent can a high-throughput, automated preprocessing pipeline maintain text quality while handling massive web-scale biomedical data?
- RQ4Can a publicly available, preprocessed corpus accelerate the development of Spanish biomedical NLP tools and models?
- RQ5What is the impact of language identification tuning on the accuracy of biomedical Spanish detection in low-resource settings?
Key findings
- The CoWeSe corpus contains 4.5GB of cleaned plain text, equivalent to 745.7 million tokens, from 1.58 million documents and 32.77 million sentences.
- The corpus was derived from 2,766 websites across diverse health domains, including hospitals, patient associations, and public health organizations.
- The preprocessing pipeline reduced the raw WARC data (905GB) to a clean, usable corpus of 4.5GB through deduplication and noise filtering.
- The corpus includes content in Spanish, Catalan, Galician, and Basque, reflecting multilingual health information in Spain.
- The corpus was released under a Creative Commons Attribution 4.0 International license, ensuring open access and reusability.
- The corpus has already been used to train domain-specific language models and generate word embeddings, demonstrating its utility in Spanish biomedical NLP.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.