[Paper Review] The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service
The paper introduces JDDC, a large-scale real-world Chinese e-commerce dialogue corpus with over 1 million multi-turn dialogues and 20 million utterances, plus extra annotations and challenge sets, and provides baseline benchmarks for retrieval-based and generative models.
Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task
Motivation & Objective
- Construct a large-scale real-scenario Chinese e-commerce conversation corpus (JDDC).
- Capture characteristics of human-human dialogue such as goal-driven interactions and long-term context dependence.
- Cover diverse dialogue types including task-oriented, chitchat, and question-answering.
- Provide extra intent information and three well-annotated challenge sets for robust evaluation.
Proposed method
- Assemble a real-scenario Chinese e-commerce corpus with over 1 million multi-turn dialogues, 20 million utterances, and 150 million words.
- Annotate extra intent information and create three challenge sets to facilitate robust evaluation.
- Benchmark baseline performance using retrieval-based and generative models on the JDDC corpus.
Experimental results
Research questions
- RQ1What baseline performance do retrieval-based models achieve on the JDDC dataset?
- RQ2What baseline performance do generative models achieve on the JDDC dataset?
- RQ3How well does JDDC reflect goal-driven behavior and long-term dependency in multi-turn dialogues?
Key findings
- The dataset contains over 1 million multi-turn dialogues, with 20 million utterances and 150 million words.
- JDDC reflects goal-driven and long-term dependency characteristics of human conversations.
- JDDC supports diverse dialogue types, including task-oriented, chitchat, and question-answering.
- Extra intent information is provided to aid analysis and modeling.
- Three well-annotated challenge sets are supplied to diversify evaluation scenarios.
- Baseline benchmarking is performed for both retrieval-based and generative models on JDDC.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.