LLaVA
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
About
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Details
LLaVA (Large Language and Vision Assistant) is an open-source multimodal model that combines a vision encoder with a large language model for visual instruction tuning. Presented as a NeurIPS 2023 Oral paper, LLaVA is designed to achieve capabilities comparable to GPT-4V, enabling it to understand and reason about images in response to natural language instructions. The model takes both image and text inputs and generates text outputs, making it suitable for a wide range of vision-language tasks. As an open-source project hosted on GitHub, LLaVA allows researchers and developers to access, modify, and deploy the model for their own use cases.