Chinese Large Model Leaderboard
Freean expert-driven benchmark for Chineses LLMs.
About Chinese Large Model Leaderboard
ReLE (Really Reliable Live Evaluation for LLM), formerly CLiB, is an open-source, continuously updated benchmark for evaluating Chinese large language models (LLMs). Hosted on GitHub by 非线智能 (NoneLinear), it covers 394+ models including commercial models like ChatGPT, Gemini, Claude, ERNIE, Qwen, and open-source models like DeepSeek, LLaMA, GLM, and more. The benchmark evaluates LLMs across seven major domains: education, medical & mental health, finance, law & administrative affairs, reasoning & math calculation, language & instruction following, and agent & tool calling, with approximately 300 sub-dimensions (e.g., dentistry, high school Chinese). It provides a comprehensive leaderboard with rankings for overall capability and specific abilities, as well as a defect library containing over 2 million entries to aid community research and model improvement. The project also offers free evaluation services for private large models and includes a technical report and media coverage from outlets like 机器之心 (Machine Heart).
Key Features
Pros & Cons
- Extensive model coverage spanning both commercial and open-source
- Detailed multi-domain and sub-dimension evaluation for nuanced insights
- Large defect library facilitates deep analysis and model refinement
- Actively maintained with frequent updates and new model additions
- Free evaluation service available for private models
- Primarily focused on Chinese language models, limiting applicability to other languages
- Project hosted on GitHub with a community-driven interface, not a polished commercial product