LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
FreeStep-by-step visual reasoning in LMMs
About LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
LlamaV-o1 is a multimodal visual reasoning model designed for step-by-step reasoning in large language models (LMMs). It is trained using a multi-step curriculum learning approach that progressively organizes tasks to facilitate incremental skill acquisition. The project also introduces a visual reasoning benchmark with over 4,000 reasoning steps across eight categories (e.g., complex visual perception and scientific reasoning) and a novel metric that evaluates reasoning quality at the granularity of individual steps, emphasizing both correctness and logical coherence. The model achieves state-of-the-art performance among open-source models, outperforming Llava-CoT with an average score of 67.3 (3.8% absolute gain) while being 5 times faster during inference scaling. The benchmark, model, and code are publicly available.
Key Features
Pros & Cons
- Outperforms existing open-source models across multiple benchmarks
- Achieves 5x faster inference scaling than Llava-CoT
- Introduces comprehensive benchmark and step-wise metric for transparent evaluation
- Open-source with publicly available code, model, and data
- Training via multi-step curriculum learning improves skill acquisition