Skip to main content
QUICK REVIEW

[Paper Review] NusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages

Samuel Cahyawijaya, Alham Fikri Aji|arXiv (Cornell University)|Jul 21, 2022
Software Engineering Research4 citations
TL;DR

NusaCrowd proposes a centralized, open, and reproducible NLP initiative for Indonesian languages, aggregating datasets via a curated datasheet registry (NusaCatalogue) and standardized dataloaders (NusaCrowd Data Hub). By enabling community-driven contributions—datasheet submissions, dataloader implementations, and private dataset releases—it addresses data scarcity and fragmentation, fostering collaboration and reproducibility in low-resource Indonesian NLP research.

ABSTRACT

At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarce and underrepresented. Many Indonesian researchers do not publish their dataset. Furthermore, the few public datasets that we have are scattered across different platforms, thus makes performing reproducible and data-centric research in Indonesian NLP even more arduous. Rising to this challenge, we initiate the first Indonesian NLP crowdsourcing effort, NusaCrowd. NusaCrowd strives to provide the largest datasheets aggregation with standardized data loading for NLP tasks in all Indonesian languages. By enabling open and centralized access to Indonesian NLP resources, we hope NusaCrowd can tackle the data scarcity problem hindering NLP progress in Indonesia and bring NLP practitioners to move towards collaboration.

Motivation & Objective

  • Address the severe scarcity and fragmentation of Indonesian NLP datasets, especially for local languages.
  • Overcome the lack of public, discoverable, and reusable datasets in Indonesian NLP research.
  • Establish a centralized, open-access platform to improve data discoverability, interoperability, and reproducibility.
  • Foster community collaboration through structured contribution pathways and a transparent scoring system.
  • Promote data sharing by incentivizing private dataset release through contribution points and recognition.

Proposed method

  • Curate and host public NLP dataset metadata (datasheets) on NusaCatalogue, a searchable web portal inspired by the datasheet framework.
  • Implement standardized, programmatic dataloader scripts for each dataset via GitHub, ensuring consistent and reproducible data access.
  • Enforce quality control through a hybrid evaluation system combining automatic checks and manual review of submissions.
  • Use a contribution point system to incentivize participation, with points awarded for datasheet submission, dataloader implementation, and private dataset release.
  • Maintain data integrity by preserving original licensing and linking directly to source repositories and publications.
  • Organize the initiative as a time-bound, community-driven movement (June 25–November 18, 2022), with weekly progress tracking and final paper submission to ACL 2023.

Experimental results

Research questions

  • RQ1How can data scarcity and fragmentation in Indonesian NLP be mitigated through centralized, community-driven curation?
  • RQ2To what extent can a standardized, open-access platform improve the discoverability and reproducibility of Indonesian NLP datasets?
  • RQ3What impact does a contribution point system have on motivating researchers to share private or underutilized NLP datasets?
  • RQ4Can a collaborative, open framework effectively unify scattered NLP resources across diverse Indonesian languages and institutions?
  • RQ5How can standardized dataloaders and metadata improve interoperability and reduce barriers to entry in Indonesian NLP research?

Key findings

  • NusaCrowd successfully established a centralized, open-access platform for Indonesian NLP datasets through NusaCatalogue and the NusaCrowd Data Hub.
  • The initiative enabled community-driven contributions, including datasheet submissions and dataloader implementations, across multiple Indonesian languages.
  • A transparent contribution point system was implemented, with points awarded for datasheet submission, dataloader development, and private dataset release.
  • The project maintained data integrity by preserving original licenses and linking directly to source repositories and publications.
  • The movement culminated in a final contribution matrix and a research paper submitted to ACL 2023, demonstrating the feasibility of large-scale, collaborative NLP data curation in low-resource settings.
  • The initiative has fostered a growing community of contributors through Slack, WhatsApp, and GitHub, signaling sustained engagement potential.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.