Skip to main content
QUICK REVIEW

[Paper Review] MedDialog: A Large-scale Medical Dialogue Dataset

Shu Chen, Zeqian Ju|arXiv (Cornell University)|Apr 7, 2020
Machine Learning in Healthcare2 references28 citations
TL;DR

This paper introduces MedDialog, a large-scale medical dialogue dataset comprising 0.3 million English and 1.1 million Chinese patient-doctor conversations, totaling 0.5 million and 4 million utterances respectively. It is the largest medical dialogue dataset to date, designed to advance research in medical dialogue systems for telemedicine applications.

ABSTRACT

Medical dialogue systems are promising in assisting in telemedicine to increase access to healthcare services, improve the quality of patient care, and reduce medical costs. To facilitate the research and development of medical dialogue systems, we build two large-scale medical dialogue datasets: MedDialog-EN and MedDialog-CN. MedDialog-EN is an English dataset containing 0.3 million conversations between patients and doctors and 0.5 million utterances. MedDialog-CN is an Chinese dataset containing 1.1 million conversations and 4 million utterances. To our best knowledge, MedDialog-(EN,CN) are the largest medical dialogue datasets to date. The dataset is available at this https URL

Motivation & Objective

  • To address the lack of large-scale, diverse medical dialogue datasets for training and evaluating medical dialogue systems.
  • To improve access to healthcare through scalable, data-driven medical dialogue systems in telemedicine.
  • To support multilingual research by providing parallel English and Chinese medical dialogue data.
  • To facilitate the development of accurate and context-aware medical conversational agents for improved patient care.

Proposed method

  • Curated and constructed MedDialog-EN and MedDialog-CN from real-world medical consultation data.
  • Collected patient-doctor conversations from telemedicine platforms, ensuring clinical relevance and diversity.
  • Annotated and structured the data into turn-by-turn dialogue formats for model training and evaluation.
  • Ensured data quality through expert review and linguistic validation to maintain medical accuracy.
  • Standardized dialogue formatting and metadata for interoperability with existing NLP and dialogue system frameworks.
  • Released the datasets publicly at a dedicated URL to enable broad research access and reproducibility.

Experimental results

Research questions

  • RQ1Can a large-scale, multilingual medical dialogue dataset improve the performance of medical dialogue systems in real-world telemedicine settings?
  • RQ2How does the scale and diversity of dialogue data impact the training of context-aware medical conversational agents?
  • RQ3To what extent can multilingual medical dialogue datasets support the development of cross-lingual medical dialogue models?
  • RQ4What are the key challenges in collecting and curating large-scale, clinically accurate medical dialogue data at scale?

Key findings

  • MedDialog-EN contains 0.3 million English conversations with 0.5 million utterances, making it one of the largest publicly available English medical dialogue datasets.
  • MedDialog-CN comprises 1.1 million Chinese conversations and 4 million utterances, significantly expanding the scale of Chinese medical dialogue data.
  • The combined datasets represent the largest medical dialogue collection to date in both English and Chinese.
  • The datasets are publicly released at a dedicated URL to support open research and reproducibility in medical dialogue systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.