EN Submit a tool

LLMEval3

LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.

Human verified · · Submit a correction

Editor's note

Best for research teams and edtech organizations that need to probe deep professional knowledge; not for those only interested in general chat ability.

Verified facts

CategoryLLMs
Human verified

What is LLMEval3

LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.

Key features of LLMEval3

  • Evaluating professional knowledge of LLMs across 13 discipline categories
  • Comparing vertical-domain knowledge depth between models
  • Selecting foundation models for knowledge-intensive applications
  • Supporting academic research on Chinese LLM capability

Good for

  • Focused on professional knowledge with systematic discipline coverage
  • About 200,000 standardized generative QA items provide ample test volume
  • Designed by a university NLP lab with strong academic rigor

Watch out

  • Focuses on knowledge QA, with limited coverage of newer abilities like tool use
  • Public page information is brief; usage details require further inquiry
  • Judging generative answers demands a robust evaluation methodology

How to use LLMEval3

  1. Visit llmeval.com to learn the LLMEval-3 evaluation design
  2. Review the coverage of 13 discipline categories and sub-disciplines
  3. Understand how the generative QA items are organized
  4. Follow the page guidance to use or cite the benchmark
  5. Compare per-discipline results across models

Who LLMEval3 is for

Difficulty: Intermediate

  • Evaluating professional knowledge of LLMs across 13 discipline categories
  • Comparing vertical-domain knowledge depth between models
  • Selecting foundation models for knowledge-intensive applications
  • Supporting academic research on Chinese LLM capability

FAQ

Who created LLMEval?

It was created by Fudan University's NLP Lab; the latest version, LLMEval-3, focuses on professional knowledge evaluation.

Which disciplines does LLMEval-3 cover?

Thirteen discipline categories including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management and art, across 50+ sub-disciplines.

How large is the LLMEval-3 question set?

About 200,000 standardized generative question-answering items.

Sources and verification

Sources: llmeval.com (opens in a new tab)
Verified: · Submit a correction →

Alternatives to LLMEval3

All in this category