DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning logo

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Free

A vision-language model that learns to think with images through reinforcement learning

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

DeepEyes is a large vision-language model trained end-to-end via reinforcement learning to integrate visual information into its reasoning processes, a capability known as 'thinking with images'. Unlike traditional approaches, DeepEyes does not require pre-collected reasoning data for cold-start supervised fine-tuning; instead, it learns to ground its reasoning in visual information through active perception, guided by a tailored data selection and reward strategy. The model achieves significant performance gains on general perception and reasoning benchmarks, and also demonstrates improvements in grounding, hallucination reduction, and mathematical reasoning tasks. Notably, it exhibits diverse thinking patterns that closely mirror human visual reasoning processes. The code is publicly available.

Key Features

End-to-end reinforcement learning without cold-start supervised fine-tuning
Active perception: strategically grounds reasoning in visual information
Tailored data selection and reward strategy
Significant performance gains on perception and reasoning benchmarks
Improvements in grounding, hallucination, and mathematical reasoning
Diverse thinking patterns that mimic human visual reasoning
Open-source code available

Pros & Cons

Pros
  • No need for pre-collected reasoning data for cold-start supervised fine-tuning
  • End-to-end reinforcement learning simplifies training pipeline
  • Achieves significant performance gains on multiple benchmark types
  • Demonstrates improvement in grounding and hallucination tasks
  • Exhibits diverse thinking patterns similar to human visual reasoning

Best For

General perception and reasoning tasksGrounding tasks that require integrating visual informationReducing hallucination in vision-language responsesMathematical reasoning involving visual elements

FAQ

How does DeepEyes train without pre-collected reasoning data?
DeepEyes uses reinforcement learning end-to-end, leveraging its own grounding capability as an intrinsic function, without requiring cold-start supervised fine-tuning.
What tasks does DeepEyes improve?
DeepEyes achieves performance gains on general perception, reasoning, grounding, hallucination, and mathematical reasoning tasks.