DreamBench++ logo

DreamBench++

Free

a benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.

FreeFree tier
Type
Open Source

About DreamBench++

DreamBench++ is a human-aligned benchmark for personalized image generation, presented at ICLR 2025. It systematically designs prompts to leverage GPT-4o for automated evaluation that aligns with human preferences through task reinforcement and self-alignment. The benchmark includes a diverse dataset of images and prompts, and evaluates 7 modern generative models on concept preservation and prompt following across multiple sub-categories (Animal, Human, Object, Style, Photorealistic, Style, Imaginative). It provides a leaderboard and detailed quality results to help advance the field.

Key Features

Human-aligned automated evaluation using GPT-4o with task reinforcement
Comprehensive dataset with diverse images and prompts
Benchmarks 7 modern generative models (e.g., DreamBooth, IP-Adapter, Emu2, BLIP-Diffusion, Textual Inversion)
Evaluates concept preservation and prompt following across multiple categories (Animal, Human, Object, Style, Photorealistic, Style, Imaginative)
Provides leaderboard and detailed quality results
Self-aligned prompt strategy for GPT-4o including reasoning instructions

Pros & Cons

Pros
  • Automated evaluation that aligns with human preferences
  • Uses advanced GPT-4o model for scoring
  • Open source code and dataset available (arXiv, Code links)
  • Comprehensive benchmark covering multiple dimensions of quality
Cons
  • Limited to evaluation using GPT-4o, which may have inherent biases
  • Only evaluates 7 models; may not cover all recent approaches
  • Requires access to GPT-4o API, which can be costly for large-scale use
  • Benchmark dataset size and diversity may not generalize to all real-world scenarios

Best For

Evaluating personalized image generation models for researchComparing model performance on concept preservation and prompt followingBenchmarking new generative methods against state-of-the-artStudying human-aligned evaluation metrics