WHOOPS!
Freea benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.
About WHOOPS!
WHOOPS! is a benchmark dataset and evaluation suite for visual commonsense reasoning. It consists of 500 synthetic images that deliberately defy commonsense expectations, created by designers using text-to-image models such as Midjourney and DALL-E. The dataset includes 10,874 human annotations and spans four tasks: image captioning, cross-modal matching, visual question answering (VQA), and a novel explanation-of-violation task where models must identify and explain why an image is unusual. Evaluation of state-of-the-art vision-language models (e.g., BLIP2, GPT-3) shows they still lag significantly behind human performance, highlighting the challenge WHOOPS! poses for advancing AI's understanding of compositionality and everyday knowledge.
Key Features
Pros & Cons
- Challenging benchmark that exposes weaknesses of current models
- Includes a novel explanation-of-violation task not present in other datasets
- High-quality human annotations with detailed explanations
- Covers a wide range of commonsense violations (social norms, everyday knowledge)
- Relatively small dataset with only 500 images
- Synthetic images may not fully capture real-world unusual scenarios
- State-of-the-art models significantly lag behind human performance