SuperCLUE
SuperCLUE is a comprehensive Chinese LLM benchmark spanning four capability quadrants and twelve abilities. It blends objective and subjective tests, covers Chinese-specific and agent skills, and updates monthly.
Human verified · · Submit a correction
Best for: Chinese LLM teams, enterprise selectors, and researchers tracking Chinese models; not for: minimal use cases needing only quick objective scoring.
Verified facts
What is SuperCLUE
SuperCLUE is a comprehensive Chinese LLM benchmark spanning four capability quadrants and twelve abilities. It blends objective and subjective tests, covers Chinese-specific and agent skills, and updates monthly.
Key features of SuperCLUE
- Evaluating a model across twelve Chinese capability dimensions
- Testing multi-turn dialogue coherence and context tracking
- Probing Chinese-specific skills like idioms, poetry, and dialects
- Assessing agent tool use and task planning
Good for
- Four quadrants, twelve abilities, objective plus subjective tests
- Deep Chinese-specific testing rather than translated exams
- Monthly leaderboard updates with detailed technical reports
Watch out
- Submission-based; not every model is scored
- Subjective grading stability requires reading the reports
- Submission terms and fees are not published
How to use SuperCLUE
- Read the technical reports on the site or GitHub
- Ensure your model can interact with the evaluation via API
- Submit the model through the official CLUEbenchmark email
- Check results and rankings on the leaderboard
- Study per-ability analysis in the reports
Who SuperCLUE is for
Difficulty: Advanced
- Evaluating a model across twelve Chinese capability dimensions
- Testing multi-turn dialogue coherence and context tracking
- Probing Chinese-specific skills like idioms, poetry, and dialects
- Assessing agent tool use and task planning
FAQ
What abilities does SuperCLUE evaluate?
Twelve abilities in four quadrants: language understanding and generation, knowledge application, professional skills including coding and agents, plus adaptation and safety, with Chinese-specific tests.
How do I get my model on the leaderboard?
Contact the organizers via the official CLUEbenchmark email, submit model information, and they run the evaluation.
How often is the leaderboard updated?
Monthly, accompanied by technical reports analyzing each model's strengths and weaknesses.
Sources and verification
Sources: www.cluebenchmarks.com (opens in a new tab)
Verified: · Submit a correction →