LLMEval
Freefocuses on understanding how these models perform in various scenarios and analyzing results from an interpretability perspective.
About LLMEval
LLMEval is a research initiative from Fudan NLP Lab that develops rigorous and fair evaluation frameworks for large language models (LLMs). It encompasses multiple benchmarks and datasets across 13+ academic disciplines, medical AI, and logical reasoning, with over 220,000 generative questions in LLMEval-Fair alone. The project emphasizes contamination-resistant evaluation, adversarial hardening, and reproducibility, and has published papers at top conferences including AAAI, EMNLP, ACL, and arXiv. LLMEval-Logic, a Chinese logical reasoning benchmark built with Z3 verification and adversarial hardening, is open source on GitHub.
Key Features
Pros & Cons
- Open-source and transparent evaluation pipeline
- Contamination-resistant design with private test sets
- High agreement with human experts (90%) in LLM-as-a-judge
- Covers a wide range of disciplines and reasoning types
- Includes physician-validated medical benchmark
- Adversarial hardening improves benchmark difficulty and discriminability