Qwen2-VL-2B|7B|72B logo

Qwen2-VL-2B|7B|72B

Free

To See the World More Clearly

FreeFree tier
Inputs: image, video, textOutputs: text
Type
Open Source
Company
Qwen Team

About Qwen2-VL-2B|7B|72B

Qwen2-VL is the latest vision-language model from the Qwen Team, built upon the Qwen2 architecture. It achieves state-of-the-art performance on visual understanding benchmarks (MathVista, DocVQA, RealWorldQA, MTVQA, etc.) and introduces capabilities such as understanding videos over 20 minutes, acting as an agent to control mobile phones and robots, and multilingual text-in-image recognition (covering most European languages, Japanese, Korean, Arabic, Vietnamese, etc.). The model is released in three sizes: Qwen2-VL-2B and 7B are open-sourced under Apache 2.0, while the 72B model is available via API. The 2B variant is optimized for mobile deployment, and all sizes support image, multi-image, and video inputs. Qwen2-VL demonstrates particular strength in document understanding and multilingual image-text comprehension, often surpassing GPT-4o and Claude 3.5-Sonnet at the 72B scale.

Key Features

State-of-the-art understanding of images at various resolutions and aspect ratios
Can understand videos longer than 20 minutes for high-quality QA and content creation
Agent capabilities: can operate mobile phones, robots, and other devices via visual environment and text instructions
Multilingual text recognition inside images (English, Chinese, most European languages, Japanese, Korean, Arabic, Vietnamese, etc.)
Three model sizes: 2B (mobile-optimized), 7B, and 72B (API-only)
Open-source Qwen2-VL-2B and 7B under Apache 2.0 license
Integrated with Hugging Face Transformers, vLLM, and other third-party frameworks
Enhanced object recognition including complex multi-object relationships and handwritten text

Pros & Cons

Pros
  • State-of-the-art performance on multiple vision-language benchmarks
  • Surpasses closed-source models like GPT-4o and Claude 3.5-Sonnet on many metrics
  • Open-source 2B and 7B models available for free under Apache 2.0 license
  • Strong document understanding and multilingual text recognition
  • Supports long video understanding (20+ minutes)
  • Agent capabilities enable integration with devices for automation
  • Small 2B model suitable for mobile deployment without sacrificing much performance
  • Active ecosystem with integrations in major frameworks (Transformers, vLLM)
Cons
  • The large 72B model is not open-source; only accessible via API
  • Model may require significant computational resources for local deployment of 7B size
  • Relatively new model with evolving documentation and community support

Best For

Document and table comprehension (DocVQA)Complex college-level problem solving (MathVista, RealWorldQA)Multilingual text-image understanding (MTVQA)General scenario question-answering from images and videosVideo-based question answering and dialog generationAutomatic operation of mobile devices and robots based on visual inputMobile deployment with the compact 2B modelText extraction from handwritten notes and multi-language documents

FAQ

What is Qwen2-VL?
Qwen2-VL is the latest vision-language model from the Qwen Team, based on Qwen2. It offers state-of-the-art image and video understanding, multilingual text recognition, and agent capabilities.
What model sizes are available?
Qwen2-VL is available in three sizes: 2B, 7B, and 72B. The 2B and 7B models are open-sourced under Apache 2.0, while the 72B model is available through an API.
Is Qwen2-VL free to use?
Yes, Qwen2-VL-2B and Qwen2-VL-7B are open-sourced under the Apache 2.0 license, making them free to use. The 72B model is available via API, which may have usage costs.
What languages does Qwen2-VL support?
Qwen2-VL supports text understanding inside images in English, Chinese, most European languages, Japanese, Korean, Arabic, Vietnamese, and more.
Can Qwen2-VL process videos?
Yes, Qwen2-VL can understand videos longer than 20 minutes and supports video-based question answering, dialog, and content creation.