Chinese Large Model Leaderboard logo

Chinese Large Model Leaderboard

Free

an expert-driven benchmark for Chineses LLMs.

FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
非线智能 (NoneLinear)

About Chinese Large Model Leaderboard

ReLE (Really Reliable Live Evaluation for LLM), formerly CLiB, is an open-source, continuously updated benchmark for evaluating Chinese large language models (LLMs). Hosted on GitHub by 非线智能 (NoneLinear), it covers 394+ models including commercial models like ChatGPT, Gemini, Claude, ERNIE, Qwen, and open-source models like DeepSeek, LLaMA, GLM, and more. The benchmark evaluates LLMs across seven major domains: education, medical & mental health, finance, law & administrative affairs, reasoning & math calculation, language & instruction following, and agent & tool calling, with approximately 300 sub-dimensions (e.g., dentistry, high school Chinese). It provides a comprehensive leaderboard with rankings for overall capability and specific abilities, as well as a defect library containing over 2 million entries to aid community research and model improvement. The project also offers free evaluation services for private large models and includes a technical report and media coverage from outlets like 机器之心 (Machine Heart).

Key Features

Covers 394+ Chinese LLMs including commercial and open-source models
Multi-dimensional evaluation across 7 domains: education, medical/mental health, finance, law/administration, reasoning/math, language/instruction, agent/tool calling
Approximately 300 sub-dimensions for granular capability assessment
Leaderboard with overall and domain-specific rankings
Over 2 million entries in a defect library for model research and improvement
Continuously updated with new models and evaluation results
Free evaluation service for private large models
Technical report and media coverage available

Pros & Cons

Pros
  • Extensive model coverage spanning both commercial and open-source
  • Detailed multi-domain and sub-dimension evaluation for nuanced insights
  • Large defect library facilitates deep analysis and model refinement
  • Actively maintained with frequent updates and new model additions
  • Free evaluation service available for private models
Cons
  • Primarily focused on Chinese language models, limiting applicability to other languages
  • Project hosted on GitHub with a community-driven interface, not a polished commercial product

Best For

Comparing capabilities of different Chinese LLMsSelecting the best model for specific application domains (e.g., education, finance)Identifying weaknesses and defects in LLMs for targeted improvementAcademic research on LLM evaluation methodologiesBenchmarking private or custom Chinese LLMs against public models

FAQ

What models are included in the benchmark?
The benchmark currently covers 394+ models, including commercial models such as ChatGPT, GPT-5.6, Google Gemini-3.1-pro, Claude-5, ERNIE-X1.1, ERNIE-5.1, Qwen3.7-max, and open-source models like DeepSeek-v4, LLaMA4, GLM-5.2, MiniMax-M3, and many others.
How is the evaluation structured?
Evaluation covers seven major domains: education, medical & mental health, finance, law & administrative affairs, reasoning & math calculation, language & instruction following, and agent & tool calling, with approximately 300 sub-dimensions. Each model is tested across these areas to produce scores and rankings.
Is the tool free to use?
Yes, the benchmark is open source and free. Additionally, the team offers free evaluation services for private large models.
How can I contact the team or request evaluation?
Contact information is provided via WeChat on the repository page. The team (非线智能 ReLE benchmark团队) handles inquiries and evaluation requests.
What is the defect library?
The defect library contains over 2 million entries documenting specific failures or weaknesses of evaluated models, helping researchers and developers analyze and improve LLM performance.