EN Submit a tool

SuperCLUE

SuperCLUE is a comprehensive Chinese LLM benchmark spanning four capability quadrants and twelve abilities. It blends objective and subjective tests, covers Chinese-specific and agent skills, and updates monthly.

Human verified · · Submit a correction

Editor's note

Best for: Chinese LLM teams, enterprise selectors, and researchers tracking Chinese models; not for: minimal use cases needing only quick objective scoring.

Verified facts

CategoryLLMs
Human verified

What is SuperCLUE

SuperCLUE is a comprehensive Chinese LLM benchmark spanning four capability quadrants and twelve abilities. It blends objective and subjective tests, covers Chinese-specific and agent skills, and updates monthly.

Key features of SuperCLUE

  • Evaluating a model across twelve Chinese capability dimensions
  • Testing multi-turn dialogue coherence and context tracking
  • Probing Chinese-specific skills like idioms, poetry, and dialects
  • Assessing agent tool use and task planning

Good for

  • Four quadrants, twelve abilities, objective plus subjective tests
  • Deep Chinese-specific testing rather than translated exams
  • Monthly leaderboard updates with detailed technical reports

Watch out

  • Submission-based; not every model is scored
  • Subjective grading stability requires reading the reports
  • Submission terms and fees are not published

How to use SuperCLUE

  1. Read the technical reports on the site or GitHub
  2. Ensure your model can interact with the evaluation via API
  3. Submit the model through the official CLUEbenchmark email
  4. Check results and rankings on the leaderboard
  5. Study per-ability analysis in the reports

Who SuperCLUE is for

Difficulty: Advanced

  • Evaluating a model across twelve Chinese capability dimensions
  • Testing multi-turn dialogue coherence and context tracking
  • Probing Chinese-specific skills like idioms, poetry, and dialects
  • Assessing agent tool use and task planning

FAQ

What abilities does SuperCLUE evaluate?

Twelve abilities in four quadrants: language understanding and generation, knowledge application, professional skills including coding and agents, plus adaptation and safety, with Chinese-specific tests.

How do I get my model on the leaderboard?

Contact the organizers via the official CLUEbenchmark email, submit model information, and they run the evaluation.

How often is the leaderboard updated?

Monthly, accompanied by technical reports analyzing each model's strengths and weaknesses.

Sources and verification

Sources: www.cluebenchmarks.com (opens in a new tab)
Verified: · Submit a correction →

Alternatives to SuperCLUE

All in this category