Qwen2.5-VL-3|7|72B logo

Qwen2.5-VL-3|7|72B

Free

The new flagship vision-language model of Qwen

FreeFree tier
Inputs: image, video, textOutputs: text
Type
Open Source
Company
Alibaba Cloud

About Qwen2.5-VL-3|7|72B

Qwen2.5-VL is a family of vision-language models from the Qwen Team, released in January 2025. Available in three sizes (3B, 7B, and 72B) under open-source licenses, it is designed to understand images, videos, and text. Key capabilities include world-wide image recognition (objects, landmarks, celebrities, etc.), acting as an agent for computer and phone use, understanding videos over an hour long, precise visual localization (bounding boxes/points), and generating structured outputs from documents, invoices, and tables. The 72B model achieves state-of-the-art performance on multiple benchmarks, especially document and diagram understanding, while the 7B model outperforms GPT-4o-mini in several tasks and the 3B model is optimized for edge AI. Models are available on Hugging Face, ModelScope, and via Qwen Chat.

Key Features

World-wide image recognition: identifies plants, animals, landmarks, IPs, products, and celebrities
Agentic capabilities: can reason and directly use tools, including computer and phone control
Long video understanding: comprehends videos over 1 hour and can pinpoint relevant segments for events
Visual localization: generates bounding boxes or points for object positions, with stable JSON output for coordinates and attributes
Structured output generation: extracts structured content from invoices, forms, tables, and other documents
Available in 3B, 7B, and 72B parameter sizes, with base and instruct variants
Competitive performance on benchmarks including college-level problems, math, document understanding, and visual agent tasks
Smaller models (7B) outperform GPT-4o-mini on multiple tasks; 3B model beats previous Qwen2-VL 7B

Pros & Cons

Pros
  • Open-source and freely available under a permissive license
  • Strong visual recognition across diverse object categories (plants, animals, landmarks, IPs)
  • Excellent document and diagram understanding, outperforming many competitors
  • Agentic capabilities enable real-world tool use without fine-tuning
  • Long video understanding with event pinpointing
  • Multiple model sizes allow deployment from edge devices to large servers
  • High performance for the model size; 7B outperforms GPT-4o-mini on several benchmarks

Best For

Image recognition and classification across a vast range of categoriesDocument and diagram understanding and analysisVideo analysis and event detection (e.g., finding specific moments in long videos)Visual agent for automating computer or phone interactions (e.g., clicking, typing)Structured data extraction from invoices, forms, tables, and receiptsObject localization and coordinate detection for spatial understandingAcademic and research tasks requiring multimodal reasoning

FAQ

What sizes of Qwen2.5-VL are available?
Qwen2.5-VL is available in three parameter sizes: 3B, 7B, and 72B, with both base and instruct variants.
Is Qwen2.5-VL open source?
Yes, both base and instruct models are open-sourced and available on Hugging Face and ModelScope.
What can Qwen2.5-VL do with images?
It can recognize a wide range of common objects (flowers, birds, fish, insects), as well as analyze texts, charts, icons, graphics, and layouts within images. It also supports visual localization (bounding boxes/points) and structured output generation.
How long of a video can Qwen2.5-VL understand?
Qwen2.5-VL can comprehend videos over 1 hour long and can capture events by pinpointing the relevant segments.
Does Qwen2.5-VL support agentic tasks?
Yes, it can act as a visual agent that reasons and dynamically directs tools, including computer use and phone use.