WHOOPS! logo

WHOOPS!

Free

a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.

FreeFree tier
Inputs: image, textOutputs: text
Type
Open Source

About WHOOPS!

WHOOPS! is a benchmark dataset and evaluation suite for visual commonsense reasoning. It consists of 500 synthetic images that deliberately defy commonsense expectations, created by designers using text-to-image models such as Midjourney and DALL-E. The dataset includes 10,874 human annotations and spans four tasks: image captioning, cross-modal matching, visual question answering (VQA), and a novel explanation-of-violation task where models must identify and explain why an image is unusual. Evaluation of state-of-the-art vision-language models (e.g., BLIP2, GPT-3) shows they still lag significantly behind human performance, highlighting the challenge WHOOPS! poses for advancing AI's understanding of compositionality and everyday knowledge.

Key Features

Synthetic images created by designers using Midjourney and DALL-E
500 images with 10,874 human annotations
Four tasks: image captioning, cross-modal matching, visual question answering, and explanation-of-violation
Novel explanation-of-violation task requiring models to generate detailed explanations of what makes an image weird
Challenging benchmark for vision-and-language models; current models lag behind human performance

Pros & Cons

Pros
  • Challenging benchmark that exposes weaknesses of current models
  • Includes a novel explanation-of-violation task not present in other datasets
  • High-quality human annotations with detailed explanations
  • Covers a wide range of commonsense violations (social norms, everyday knowledge)
Cons
  • Relatively small dataset with only 500 images
  • Synthetic images may not fully capture real-world unusual scenarios
  • State-of-the-art models significantly lag behind human performance

Best For

Evaluating visual commonsense reasoning in AI modelsTesting vision-language models on compositionality and everyday knowledgeResearch on understanding unusual or uncanny imagesBenchmarking progress in multimodal AI systems

FAQ

What is WHOOPS!?
WHOOPS! is a benchmark dataset of 500 synthetic commonsense-defying images designed to test AI models' visual commonsense reasoning abilities.
What tasks does WHOOPS! include?
It includes image captioning, cross-modal matching, visual question answering, and a novel explanation-of-violation task where models must identify and explain why an image is unusual.
How were the images created?
The images were created by designers using text-to-image models like Midjourney and DALL-E to generate images that violate commonsense expectations.
How do AI models perform on WHOOPS!?
State-of-the-art models such as BLIP2 and GPT-3 still lag behind humans. The best end-to-end fine-tuned BLIP2 achieves 73% on identification, and the oracle explanation model achieves 68% vs. human performance of 95%.