LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs logo

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Free

Step-by-step visual reasoning in LMMs

FreeFree tier
Inputs: image, textOutputs: text
Type
Open Source

About LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

LlamaV-o1 is a multimodal visual reasoning model designed for step-by-step reasoning in large language models (LMMs). It is trained using a multi-step curriculum learning approach that progressively organizes tasks to facilitate incremental skill acquisition. The project also introduces a visual reasoning benchmark with over 4,000 reasoning steps across eight categories (e.g., complex visual perception and scientific reasoning) and a novel metric that evaluates reasoning quality at the granularity of individual steps, emphasizing both correctness and logical coherence. The model achieves state-of-the-art performance among open-source models, outperforming Llava-CoT with an average score of 67.3 (3.8% absolute gain) while being 5 times faster during inference scaling. The benchmark, model, and code are publicly available.

Key Features

Multi-step curriculum learning training paradigm
Visual reasoning benchmark with 8 categories and over 4,000 reasoning steps
Novel step-wise evaluation metric for correctness and logical coherence
5x faster inference scaling compared to Llava-CoT
Publicly available code, model, and benchmark
Designed for interpretable multi-step visual reasoning

Pros & Cons

Pros
  • Outperforms existing open-source models across multiple benchmarks
  • Achieves 5x faster inference scaling than Llava-CoT
  • Introduces comprehensive benchmark and step-wise metric for transparent evaluation
  • Open-source with publicly available code, model, and data
  • Training via multi-step curriculum learning improves skill acquisition

Best For

Complex visual perception tasksScientific reasoning in visual contextsMulti-step problem solving requiring sequential understandingEvaluating and benchmarking step-by-step reasoning accuracy