DreamBench++
Freea benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.
FreeFree tier
About DreamBench++
DreamBench++ is a human-aligned benchmark for personalized image generation, presented at ICLR 2025. It systematically designs prompts to leverage GPT-4o for automated evaluation that aligns with human preferences through task reinforcement and self-alignment. The benchmark includes a diverse dataset of images and prompts, and evaluates 7 modern generative models on concept preservation and prompt following across multiple sub-categories (Animal, Human, Object, Style, Photorealistic, Style, Imaginative). It provides a leaderboard and detailed quality results to help advance the field.
Key Features
Human-aligned automated evaluation using GPT-4o with task reinforcement
Comprehensive dataset with diverse images and prompts
Benchmarks 7 modern generative models (e.g., DreamBooth, IP-Adapter, Emu2, BLIP-Diffusion, Textual Inversion)
Evaluates concept preservation and prompt following across multiple categories (Animal, Human, Object, Style, Photorealistic, Style, Imaginative)
Provides leaderboard and detailed quality results
Self-aligned prompt strategy for GPT-4o including reasoning instructions
Pros & Cons
Pros
- Automated evaluation that aligns with human preferences
- Uses advanced GPT-4o model for scoring
- Open source code and dataset available (arXiv, Code links)
- Comprehensive benchmark covering multiple dimensions of quality
Cons
- Limited to evaluation using GPT-4o, which may have inherent biases
- Only evaluates 7 models; may not cover all recent approaches
- Requires access to GPT-4o API, which can be costly for large-scale use
- Benchmark dataset size and diversity may not generalize to all real-world scenarios
Best For
Evaluating personalized image generation models for researchComparing model performance on concept preservation and prompt followingBenchmarking new generative methods against state-of-the-artStudying human-aligned evaluation metrics