C-Eval
C-Eval is a Chinese evaluation suite for foundation models from SJTU, Tsinghua, and Edinburgh. It offers 13,948 multiple-choice questions across 52 subjects and four difficulty levels, with zero-shot and few-shot testing.
Human verified · · Submit a correction
Best for: R&D teams, researchers, and decision-makers needing standardized Chinese model measurement; not for: evaluating long-form generation or conversational experience.
Verified facts
What is C-Eval
C-Eval is a Chinese evaluation suite for foundation models from SJTU, Tsinghua, and Edinburgh. It offers 13,948 multiple-choice questions across 52 subjects and four difficulty levels, with zero-shot and few-shot testing.
Key features of C-Eval
- Benchmarking a model's Chinese multi-discipline knowledge
- Testing generalization in zero-shot or few-shot modes
- Comparing models across subjects and difficulty tiers
- Assessing domain knowledge for education, finance, or healthcare
Good for
- 52 subjects and four difficulty tiers for broad, granular coverage
- Objective multiple-choice scoring with strong comparability
- Strong academic backing, community recognition, free open data
Watch out
- Measures only multiple-choice knowledge and reasoning
- Cannot gauge long-form generation, dialogue, or instruction following
- Scores may be inflated by benchmark training; combine with other suites
How to use C-Eval
- Load or download the ceval-exam dataset from Hugging Face
- Choose zero-shot or few-shot evaluation mode
- Load the target model and tokenizer for inference
- Run the model over all questions and collect outputs
- Score accuracy by subject and difficulty level
Who C-Eval is for
Difficulty: Advanced
- Benchmarking a model's Chinese multi-discipline knowledge
- Testing generalization in zero-shot or few-shot modes
- Comparing models across subjects and difficulty tiers
- Assessing domain knowledge for education, finance, or healthcare
FAQ
How many questions and subjects does C-Eval have?
13,948 multiple-choice questions across 52 subjects spanning STEM, social sciences, and humanities, in four difficulty tiers.
What is the difference between zero-shot and few-shot?
Zero-shot gives no examples; few-shot provides a few. They measure cold performance versus guided performance respectively.
Is C-Eval free?
Yes. The dataset is publicly available on Hugging Face; your only cost is evaluation compute.
Sources and verification
Sources: cevalbenchmark.com (opens in a new tab)
Verified: · Submit a correction →