MOSAIC: Granular Instruction Following Evaluation (2026) logo

MOSAIC: Granular Instruction Following Evaluation (2026)

Free

Modular benchmark with up to 20 application-oriented generation constraints per prompt; finds compliance degrades with constraint count and position (primacy/recency bias) — exposes multi-instruction conflict effects

FreeFree tier
Type
Open Source

About MOSAIC: Granular Instruction Following Evaluation (2026)

MOSAIC (MOdular Synthetic Assessment of Instruction Compliance) is a benchmark framework introduced in a paper accepted to EACL 2026 for granular evaluation of large language model instruction-following abilities. It uses a dynamically generated dataset with up to 20 application-oriented generation constraints per prompt to isolate compliance from task success. Evaluated on five LLMs from different families, MOSAIC reveals that compliance is not monolithic but varies with constraint type, quantity, and position. It uncovers model-specific weaknesses, synergistic and conflicting interactions between instructions, and distinct positional biases (primacy and recency effects), enabling detailed diagnosis of model failures and supporting development of more reliable LLMs for complex instruction scenarios.

Key Features

Modular framework for synthetic assessment of instruction compliance
Dynamically generated dataset with up to 20 constraints per prompt
Granular evaluation independent of task success
Identifies model-specific weaknesses and instruction interaction effects
Detects primacy and recency position biases
Evaluated across five LLMs from different families

Pros & Cons

Pros
  • Granular analysis of instruction compliance per constraint type, quantity, and position
  • Reveals synergistic and conflicting interactions between multiple instructions
  • Highlights model-specific weaknesses not visible in standard benchmarks
  • Helps improve reliability of LLMs in real-world applications
Cons
  • Limited to evaluation of five LLMs in the initial study
  • Focuses on synthetic constraints; real-world generalizability may need further validation
  • Requires access to LLM outputs for evaluation (no integrated tooling provided)
  • Benchmark is research-oriented, not a ready-to-use product

Best For

Evaluating and comparing LLM instruction-following capabilitiesDiagnosing model failures in complex multi-constraint scenariosDeveloping more reliable LLMs for systems demanding strict adherence to instructionsResearch on instruction compliance and positional biases

FAQ

What does MOSAIC stand for?
MOSAIC stands for MOdular Synthetic Assessment of Instruction Compliance.
How many constraints can be included in a single prompt?
MOSAIC supports up to 20 application-oriented generation constraints per prompt.
Which biases does MOSAIC detect?
MOSAIC identifies primacy and recency biases, where constraints at the beginning or end of a prompt are more likely to be followed.
How many LLMs were evaluated in the paper?
Five LLMs from different model families were evaluated using MOSAIC.
Where was MOSAIC accepted?
The paper is accepted to EACL 2026 (European Chapter of the Association for Computational Linguistics).