DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
FreeA vision-language model that learns to think with images through reinforcement learning
About DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
DeepEyes is a large vision-language model trained end-to-end via reinforcement learning to integrate visual information into its reasoning processes, a capability known as 'thinking with images'. Unlike traditional approaches, DeepEyes does not require pre-collected reasoning data for cold-start supervised fine-tuning; instead, it learns to ground its reasoning in visual information through active perception, guided by a tailored data selection and reward strategy. The model achieves significant performance gains on general perception and reasoning benchmarks, and also demonstrates improvements in grounding, hallucination reduction, and mathematical reasoning tasks. Notably, it exhibits diverse thinking patterns that closely mirror human visual reasoning processes. The code is publicly available.
Key Features
Pros & Cons
- No need for pre-collected reasoning data for cold-start supervised fine-tuning
- End-to-end reinforcement learning simplifies training pipeline
- Achieves significant performance gains on multiple benchmark types
- Demonstrates improvement in grounding and hallucination tasks
- Exhibits diverse thinking patterns similar to human visual reasoning