[Paper Review] KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension
KorQuAD1.0 introduces a large-scale Korean extractive QA dataset with 70,000+ human-generated question–answer pairs on Korean Wikipedia to advance MRC research.
Machine Reading Comprehension (MRC) is a task that requires machine to understand natural language and answer questions by reading a document. It is the core of automatic response technology such as chatbots and automatized customer supporting systems. We present Korean Question Answering Dataset(KorQuAD), a large-scale Korean dataset for extractive machine reading comprehension task. It consists of 70,000+ human generated question-answer pairs on Korean Wikipedia articles. We release KorQuAD1.0 and launch a challenge at https://KorQuAD.github.io to encourage the development of multilingual natural language processing research.
Motivation & Objective
- Motivate and enable robust machine reading comprehension research in Korean.
- Provide a large-scale, human-generated extractive QA dataset for Korean to support evaluation and model development.
- Promote multilingual NLP research by releasing KorQuAD1.0 and hosting a challenge.
Proposed method
- Construct a large-scale Korean QA dataset with 70,000+ human-generated question–answer pairs.
- Format the task as extractive machine reading comprehension on Korean Wikipedia articles.
- Release KorQuAD1.0 publicly and launch a challenge to foster NLP research.
Experimental results
Research questions
- RQ1How well can models perform extractive QA on Korean Wikipedia passages?
- RQ2Can the dataset facilitate advancements in multilingual NLP research and evaluation for Korean?
- RQ3What benchmarks and baselines emerge from KorQuAD1.0 as a resource for MRC in Korean?
Key findings
- The dataset contains 70,000+ human-generated question–answer pairs on Korean Wikipedia articles.
- KorQuAD1.0 is released to support and motivate multilingual NLP research in QA tasks.
- A challenge following the dataset release is launched to encourage development of MRC models for Korean.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.