InternVL-2|6|14|26 logo

InternVL-2|6|14|26

Free

Scaling up vision foundation models for generic visual-linguistic tasks

FreeFree tier
Inputs: image, textOutputs: text
Type
Open Source
Company
OpenGVLab

About InternVL-2|6|14|26

InternVL is a series of vision-language foundation models developed by OpenGVLab, scaling up vision foundation models and aligning for generic visual-linguistic tasks. The collection includes models of various sizes (6B, 14B, 19B, 34B, 40B) for tasks such as image feature extraction, visual question answering, image-text-to-text generation, and OCR. InternVL1.0 was presented at CVPR 2024 as an Oral paper, and subsequent versions (InternVL1.5, InternVL2.0, InternVL2.5, etc.) offer improved performance and Chinese language support.

Key Features

Vision-language alignment for generic tasks
Multiple model sizes: 6B, 14B, 19B, 34B, 40B parameters
Image feature extraction and visual question answering
Image-text-to-text generation with OCR support (Chinese & English)
State-of-the-art performance, CVPR 2024 Oral paper
Open-source and free to use

Pros & Cons

Pros
  • State-of-the-art results on vision-language benchmarks
  • Open-source and freely available for research and development
  • Multiple model sizes to suit different computational budgets
  • Strong OCR capabilities for multilingual text extraction
  • Published at top venues (CVPR 2024 Oral)
Cons
  • Larger model variants (40B) require substantial GPU memory and compute
  • Documentation and usage examples may be scattered across collections

Best For

Visual question answeringImage captioning and descriptionOCR and document understandingMultimodal reasoning and retrieval

FAQ

What is InternVL?
InternVL is a family of vision-language foundation models from OpenGVLab, designed for generic visual-linguistic tasks like image understanding, VQA, and OCR.
What model sizes are available?
InternVL offers variants with 6B, 14B, 19B, 34B, and 40B parameters, allowing users to choose based on performance and resource requirements.
Is InternVL open source?
Yes, InternVL models are open-source and freely available on Hugging Face under the OpenGVLab collection.
What tasks can InternVL perform?
InternVL supports image feature extraction, visual question answering, image-text-to-text generation, and OCR for both Chinese and English.