ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2023
Year
An enhanced version of the LLaVA model that incorporates a CLIP-ViT-L-336px with an MLP projection and academic-task-oriented VQA data to set new benchmarks in large multimodal models (LMM) research.
LLaVA 1.5 represents a significant step forward in large multimodal models (LMMs), which aim to integrate visual and textual understanding. By upgrading the vision encoder to CLIP-ViT-L-336px and adding an MLP projection layer, the model achieves a more nuanced alignment between image features and language. The use of academic-task-oriented VQA data further tailors the model to complex reasoning tasks, setting new benchmarks in the field. This work is particularly relevant for AI practitioners seeking to build or improve multimodal systems, as it demonstrates that targeted architectural changes and data selection can yield substantial performance gains without requiring massive model scaling.
While the abstract does not provide specific numerical metrics, it states that LLaVA 1.5 sets new benchmarks in LMM research. This implies superior performance on standard multimodal evaluation suites (e.g., VQA v2, GQA, VizWiz, etc.) compared to prior LLaVA versions and many contemporary models. The improvements are attributed to the combination of higher-resolution vision encoding, MLP projection, and targeted VQA data.
LLaVA 1.5's impact lies in its demonstration that careful architectural choices and data curation can push the state of the art in multimodal AI without requiring enormous computational resources. This encourages further research into efficient vision-language alignment and task-specific fine-tuning. For practitioners, it provides a strong baseline for building multimodal assistants, image captioning systems, and visual reasoning tools. The model's success also highlights the importance of high-resolution visual features and learned projections, which may become standard components in future LMMs.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba