Qwen2.5-VL-3|7|72B
FreeThe new flagship vision-language model of Qwen
About Qwen2.5-VL-3|7|72B
Qwen2.5-VL is a family of vision-language models from the Qwen Team, released in January 2025. Available in three sizes (3B, 7B, and 72B) under open-source licenses, it is designed to understand images, videos, and text. Key capabilities include world-wide image recognition (objects, landmarks, celebrities, etc.), acting as an agent for computer and phone use, understanding videos over an hour long, precise visual localization (bounding boxes/points), and generating structured outputs from documents, invoices, and tables. The 72B model achieves state-of-the-art performance on multiple benchmarks, especially document and diagram understanding, while the 7B model outperforms GPT-4o-mini in several tasks and the 3B model is optimized for edge AI. Models are available on Hugging Face, ModelScope, and via Qwen Chat.
Key Features
Pros & Cons
- Open-source and freely available under a permissive license
- Strong visual recognition across diverse object categories (plants, animals, landmarks, IPs)
- Excellent document and diagram understanding, outperforming many competitors
- Agentic capabilities enable real-world tool use without fine-tuning
- Long video understanding with event pinpointing
- Multiple model sizes allow deployment from edge devices to large servers
- High performance for the model size; 7B outperforms GPT-4o-mini on several benchmarks