EN Submit a tool
Back to filtered results

LLMEval3

LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.

Verification level not recorded · · Submit a correction

Editor's note

Best for research teams and edtech organizations that need to probe deep professional knowledge; not for those only interested in general chat ability.

Decision facts

“Not verified” means evidence is insufficient, not that the capability is absent.

CategoryLLMs
Evidence statusVerification level not recorded
PlatformsNot verified
AvailabilityAvailable
Chinese UINot verified
Mainland ChinaNot verified
Commercial useNot verified

What is LLMEval3

LLMEval is an LLM benchmark from Fudan University's NLP Lab; the latest LLMEval-3 focuses on professional knowledge across 13 discipline categories, 50+ sub-disciplines and about 200,000 generative QA items.

Key features of LLMEval3

  • Evaluating professional knowledge of LLMs across 13 discipline categories
  • Comparing vertical-domain knowledge depth between models
  • Selecting foundation models for knowledge-intensive applications
  • Supporting academic research on Chinese LLM capability

Good for

  • Focused on professional knowledge with systematic discipline coverage
  • About 200,000 standardized generative QA items provide ample test volume
  • Designed by a university NLP lab with strong academic rigor

Watch out

  • Focuses on knowledge QA, with limited coverage of newer abilities like tool use
  • Public page information is brief; usage details require further inquiry
  • Judging generative answers demands a robust evaluation methodology

How to use LLMEval3

  1. Visit llmeval.com to learn the LLMEval-3 evaluation design
  2. Review the coverage of 13 discipline categories and sub-disciplines
  3. Understand how the generative QA items are organized
  4. Follow the page guidance to use or cite the benchmark
  5. Compare per-discipline results across models

Who LLMEval3 is for

Difficulty: Intermediate

  • Evaluating professional knowledge of LLMs across 13 discipline categories
  • Comparing vertical-domain knowledge depth between models
  • Selecting foundation models for knowledge-intensive applications
  • Supporting academic research on Chinese LLM capability

FAQ

Who created LLMEval?

It was created by Fudan University's NLP Lab; the latest version, LLMEval-3, focuses on professional knowledge evaluation.

Which disciplines does LLMEval-3 cover?

Thirteen discipline categories including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management and art, across 50+ sub-disciplines.

How large is the LLMEval-3 question set?

About 200,000 standardized generative question-answering items.

Sources and verification

Evidence status: Verification level not recorded

Sources: llmeval.com (opens in a new tab)
Content reviewed: · Submit a correction →

Alternatives to LLMEval3

All in this category