MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes logo

MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Free

First public benchmark for detecting and correcting medical errors in clinical notes

FreeFree tier
Type
Open Source

About MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

MEDEC is the first publicly available benchmark for medical error detection and correction in clinical notes. It covers five error types: Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism. The dataset comprises 3,848 clinical texts, including 488 real clinical notes from three US hospital systems that were not previously seen by any large language model. MEDEC was used in the MEDIQA-CORR shared task and has been evaluated on recent LLMs such as o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash. A comparative study with medical doctors shows that while LLMs perform well, they are still outperformed by human experts, highlighting the benchmark's challenge and importance for improving clinical text validation.

Key Features

Covers five error types: Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism
Dataset of 3,848 clinical texts including 488 real clinical notes from three US hospital systems
Clinical notes not previously seen by any LLM
Used in the MEDIQA-CORR shared task with 17 participating systems
Evaluates recent LLMs (GPT-4, Claude 3.5 Sonnet, Gemini 2.0 Flash) and compares with medical doctors
Provides error detection and correction tasks requiring medical knowledge and reasoning

Pros & Cons

Pros
  • First publicly available benchmark for medical error detection and correction in clinical notes
  • Uses real clinical notes from multiple hospital systems, ensuring ecological validity
  • Covers a diverse set of clinically relevant error types
  • Includes human expert performance comparison, providing a strong baseline
  • Challenging benchmark that highlights gaps between LLMs and human doctors
Cons
  • LLMs still significantly underperform compared to medical doctors in error detection and correction
  • Limited to five predefined error categories, may not capture all clinical errors
  • Evaluation metrics may have limitations, as discussed in the paper
  • Dataset size (3,848 texts) may be insufficient for some training purposes

Best For

Evaluating large language models for medical text validation and error correctionTraining and fine-tuning models to detect and correct errors in clinical notesMedical NLP research to improve clinical decision support systemsBenchmarking progress in AI-assisted clinical documentation review

FAQ

What types of medical errors does MEDEC cover?
MEDEC covers five error types: Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism.
How many clinical texts are in the MEDEC dataset?
The dataset contains 3,848 clinical texts, including 488 clinical notes from three US hospital systems that were not previously seen by any LLM.
Which LLMs were evaluated on MEDEC?
Recent LLMs such as o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash were evaluated. Their performance was compared to that of medical doctors.
Is MEDEC freely available?
Yes, the dataset is publicly available. The arXiv paper provides a link to the dataset and the code for reproducing the experiments.