Skip to main content
QUICK REVIEW

[Paper Review] Measuring Massive Multitask Chinese Understanding

Hui Zeng|arXiv (Cornell University)|Apr 25, 2023
Radiomics and Machine Learning in Medical Imaging11 citations
TL;DR

The paper proposes a multitask test to assess large Chinese language models across medicine, law, psychology, and education, reporting zero-shot performance across 4 domains and subtasks.

ABSTRACT

The development of large-scale Chinese language models is flourishing, yet there is a lack of corresponding capability assessments. Therefore, we propose a test to measure the multitask accuracy of large Chinese language models. This test encompasses four major domains, including medicine, law, psychology, and education, with 15 subtasks in medicine and 8 subtasks in education. We found that the best-performing models in the zero-shot setting outperformed the worst-performing models by nearly 18.6 percentage points on average. Across the four major domains, the highest average zero-shot accuracy of all models is 0.512. In the subdomains, only the GPT-3.5-turbo model achieved a zero-shot accuracy of 0.693 in clinical medicine, which was the highest accuracy among all models across all subtasks. All models performed poorly in the legal domain, with the highest zero-shot accuracy reaching only 0.239. By comprehensively evaluating the breadth and depth of knowledge across multiple disciplines, this test can more accurately identify the shortcomings of the models.

Motivation & Objective

  • Motivate the need for comprehensive capability assessments for large Chinese language models.
  • Introduce a multitask evaluation test spanning four domains and multiple subtasks.
  • Provide zero-shot and domain-level performance insights to identify model shortcomings.

Proposed method

  • Define four domain areas (medicine, law, psychology, education) and enumerate 15 subtasks in medicine and 8 subtasks in education.
  • Evaluate large Chinese language models in a zero-shot setting on all subtasks.
  • Compare model performances to identify domain-wide and subdomain performance patterns.

Experimental results

Research questions

  • RQ1What is the zero-shot performance of large Chinese language models across four major domains?
  • RQ2Which domains or subtasks reveal the strongest or weakest model capabilities in zero-shot settings?
  • RQ3How does the best zero-shot performance compare to the worst across models and domains?

Key findings

  • The best zero-shot models outperform the worst by about 18.6 percentage points on average.
  • Across four domains, the highest average zero-shot accuracy among all models is 0.512.
  • In subdomains, GPT-3.5-turbo achieves 0.693 zero-shot accuracy in clinical medicine—the highest among all subtasks.
  • All models perform poorly in the legal domain, with the top zero-shot accuracy only 0.239.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.