[Paper Review] Measuring Massive Multitask Chinese Understanding
The paper proposes a multitask test to assess large Chinese language models across medicine, law, psychology, and education, reporting zero-shot performance across 4 domains and subtasks.
The development of large-scale Chinese language models is flourishing, yet there is a lack of corresponding capability assessments. Therefore, we propose a test to measure the multitask accuracy of large Chinese language models. This test encompasses four major domains, including medicine, law, psychology, and education, with 15 subtasks in medicine and 8 subtasks in education. We found that the best-performing models in the zero-shot setting outperformed the worst-performing models by nearly 18.6 percentage points on average. Across the four major domains, the highest average zero-shot accuracy of all models is 0.512. In the subdomains, only the GPT-3.5-turbo model achieved a zero-shot accuracy of 0.693 in clinical medicine, which was the highest accuracy among all models across all subtasks. All models performed poorly in the legal domain, with the highest zero-shot accuracy reaching only 0.239. By comprehensively evaluating the breadth and depth of knowledge across multiple disciplines, this test can more accurately identify the shortcomings of the models.
Motivation & Objective
- Motivate the need for comprehensive capability assessments for large Chinese language models.
- Introduce a multitask evaluation test spanning four domains and multiple subtasks.
- Provide zero-shot and domain-level performance insights to identify model shortcomings.
Proposed method
- Define four domain areas (medicine, law, psychology, education) and enumerate 15 subtasks in medicine and 8 subtasks in education.
- Evaluate large Chinese language models in a zero-shot setting on all subtasks.
- Compare model performances to identify domain-wide and subdomain performance patterns.
Experimental results
Research questions
- RQ1What is the zero-shot performance of large Chinese language models across four major domains?
- RQ2Which domains or subtasks reveal the strongest or weakest model capabilities in zero-shot settings?
- RQ3How does the best zero-shot performance compare to the worst across models and domains?
Key findings
- The best zero-shot models outperform the worst by about 18.6 percentage points on average.
- Across four domains, the highest average zero-shot accuracy among all models is 0.512.
- In subdomains, GPT-3.5-turbo achieves 0.693 zero-shot accuracy in clinical medicine—the highest among all subtasks.
- All models perform poorly in the legal domain, with the top zero-shot accuracy only 0.239.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.