LLMEval3
LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.
Human verified · · Submit a correction
Best for research teams and edtech organizations that need to probe deep professional knowledge; not for those only interested in general chat ability.
Verified facts
What is LLMEval3
LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.
Key features of LLMEval3
- Evaluating professional knowledge of LLMs across 13 discipline categories
- Comparing vertical-domain knowledge depth between models
- Selecting foundation models for knowledge-intensive applications
- Supporting academic research on Chinese LLM capability
Good for
- Focused on professional knowledge with systematic discipline coverage
- About 200,000 standardized generative QA items provide ample test volume
- Designed by a university NLP lab with strong academic rigor
Watch out
- Focuses on knowledge QA, with limited coverage of newer abilities like tool use
- Public page information is brief; usage details require further inquiry
- Judging generative answers demands a robust evaluation methodology
How to use LLMEval3
- Visit llmeval.com to learn the LLMEval-3 evaluation design
- Review the coverage of 13 discipline categories and sub-disciplines
- Understand how the generative QA items are organized
- Follow the page guidance to use or cite the benchmark
- Compare per-discipline results across models
Who LLMEval3 is for
Difficulty: Intermediate
- Evaluating professional knowledge of LLMs across 13 discipline categories
- Comparing vertical-domain knowledge depth between models
- Selecting foundation models for knowledge-intensive applications
- Supporting academic research on Chinese LLM capability
FAQ
Who created LLMEval?
It was created by Fudan University's NLP Lab; the latest version, LLMEval-3, focuses on professional knowledge evaluation.
Which disciplines does LLMEval-3 cover?
Thirteen discipline categories including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management and art, across 50+ sub-disciplines.
How large is the LLMEval-3 question set?
About 200,000 standardized generative question-answering items.
Sources and verification
Sources: llmeval.com (opens in a new tab)
Verified: · Submit a correction →