LLMEval logo

LLMEval

Free

focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability perspective.

FreeFree tier
Type
Open Source
Company
Fudan NLP Lab

About LLMEval

LLMEval is a research initiative from Fudan NLP Lab that develops rigorous and fair evaluation frameworks for large language models (LLMs). It encompasses multiple benchmarks and datasets across 13+ academic disciplines, medical AI, and logical reasoning, with over 220,000 generative questions in LLMEval-Fair alone. The project emphasizes contamination-resistant evaluation, adversarial hardening, and reproducibility, and has published papers at top conferences including AAAI, EMNLP, ACL, and arXiv. LLMEval-Logic, a Chinese logical reasoning benchmark built with Z3 verification and adversarial hardening, is open source on GitHub.

Key Features

Covers 13+ academic disciplines with graduate-level questions
Includes medical AI benchmark (LLMEval-Med) with physician validation
Contamination-resistant data curation and anti-cheating architecture
LLM-as-a-judge pipeline achieving 90% agreement with human experts
Open-source logical reasoning benchmark (LLMEval-Logic) with Z3 verification and adversarial hardening
30-month longitudinal study with 59 LLMs benchmarked
Private held-out test set for contamination resistance

Pros & Cons

Pros
  • Open-source and transparent evaluation pipeline
  • Contamination-resistant design with private test sets
  • High agreement with human experts (90%) in LLM-as-a-judge
  • Covers a wide range of disciplines and reasoning types
  • Includes physician-validated medical benchmark
  • Adversarial hardening improves benchmark difficulty and discriminability

Best For

Evaluating logical reasoning capabilities of LLMs with adversarial hardeningFair and robust benchmarking of LLMs across multiple academic disciplinesMedical AI evaluation for clinical reasoning tasksLongitudinal performance tracking of LLMs over timeResearch on data contamination and evaluation fairness

FAQ

What is LLMEval?
LLMEval is a research initiative from Fudan NLP Lab focused on comprehensive evaluation of large language models across 13+ academic disciplines, medical AI, and logical reasoning.
Is LLMEval open source?
Yes, LLMEval-Logic is open source and available on GitHub. Other components like LLMEval-Fair and LLMEval-Med also have publicly released datasets and code.
What datasets are included in LLMEval?
LLMEval includes LLMEval-Fair (220K generative questions across 13 disciplines), LLMEval-Med (clinical physician-validated benchmark), and LLMEval-Logic (Chinese logical reasoning benchmark with base and hard splits).
How does LLMEval ensure evaluation fairness?
LLMEval uses a contamination-resistant data curation pipeline, dynamic unseen test sampling for each evaluation run, anti-cheating architecture, and a calibrated LLM-as-a-judge process with 90% agreement to human experts.