LLMEval3
LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.
Verification level not recorded · · Submit a correction
Best for research teams and edtech organizations that need to probe deep professional knowledge; not for those only interested in general chat ability.
Decision facts
“Not verified” means evidence is insufficient, not that the capability is absent.
What is LLMEval3
LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.
Key features of LLMEval3
- Evaluating professional knowledge of LLMs across 13 discipline categories
- Comparing vertical-domain knowledge depth between models
- Selecting foundation models for knowledge-intensive applications
- Supporting academic research on Chinese LLM capability
Good for
- Focused on professional knowledge with systematic discipline coverage
- About 200,000 standardized generative QA items provide ample test volume
- Designed by a university NLP lab with strong academic rigor
Watch out
- Focuses on knowledge QA, with limited coverage of newer abilities like tool use
- Public page information is brief; usage details require further inquiry
- Judging generative answers demands a robust evaluation methodology
How to use LLMEval3
- Visit llmeval.com to learn the LLMEval-3 evaluation design
- Review the coverage of 13 discipline categories and sub-disciplines
- Understand how the generative QA items are organized
- Follow the page guidance to use or cite the benchmark
- Compare per-discipline results across models
Who LLMEval3 is for
Difficulty: Intermediate
- Evaluating professional knowledge of LLMs across 13 discipline categories
- Comparing vertical-domain knowledge depth between models
- Selecting foundation models for knowledge-intensive applications
- Supporting academic research on Chinese LLM capability
FAQ
Who created LLMEval?
It was created by Fudan University's NLP Lab; the latest version, LLMEval-3, focuses on professional knowledge evaluation.
Which disciplines does LLMEval-3 cover?
Thirteen discipline categories including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management and art, across 50+ sub-disciplines.
How large is the LLMEval-3 question set?
About 200,000 standardized generative question-answering items.
Sources and verification
Evidence status: Verification level not recorded
Sources: llmeval.com (opens in a new tab)
Content reviewed: · Submit a correction →