Benchmarks · Language & multilingual

CMMLU

also: Chinese Massive Multitask Language Understanding

How much a model knows across a broad span of subjects when tested natively in Chinese, including subjects specific to Chinese culture, history and civil-service-style knowledge that an English test would not cover at all.

MBZUAI, LibrAI, Shanghai Jiao Tong University, Microsoft Research Asia & University of MelbourneReleased 15 June 2023Live

CMMLU extends the MMLU model — a broad, multi-subject, multiple-choice test of general knowledge — into Chinese, but rather than translating the original American test bank, its authors built questions from scratch around subjects that only make sense in a Chinese context: driving-test rules, Chinese food culture, elementary Chinese language instruction, and civil-service-style general knowledge, alongside the sciences, humanities and professional subjects a Western equivalent would also cover. Researchers from MBZUAI, LibrAI, Shanghai Jiao Tong University, Microsoft Research Asia and the University of Melbourne assembled 11,528 questions across 67 subjects spanning elementary through advanced professional difficulty.

In the paper’s own evaluation of 18 models, GPT-4 led with 71.0% average accuracy in a five-shot setting, well ahead of other systems tested — evidence, the authors argued, that even strong general-purpose models had real room to improve on knowledge specific to Chinese language and culture rather than knowledge translatable from English. Most other models fell short of 50% despite in-context examples and chain-of-thought prompting, against a 25% random baseline.

CMMLU sits alongside C-Eval as one of two standard Chinese-language knowledge benchmarks that emerged in the same few months of 2023, and the two are usually reported together in Chinese model releases: C-Eval leans more heavily on exam-style academic material, while CMMLU deliberately reaches for content with no English equivalent at all.

The set

11,528 multiple-choice questions across 67 subjects, from elementary to advanced professional level, including China-specific topics such as Chinese driving rules, food culture and elementary Chinese; a development split of 5 questions per subject and a test split of over 100.

Example

A high-school biology question, posed in Chinese: two cell types in the same species make secretory proteins with identical amino acids but different sequence order — why? Options point to tRNA, codon, mRNA and ribosome differences; the correct answer is differing mRNA base sequences.github.com

Where it stands

Widely reported alongside MMLU and C-Eval in Chinese model releases as of 2024; the authors note most models still fall well short of the roughly 90% ceiling implied by human performance.

How the top score changed hands

  1. June 2023GPT-471.0% average accuracyFive-shot setting; the highest score reported in the original paper's evaluation.

More language & multilingual benchmarks