Preprint2023
LLaVA 1.5
Unknown
LLaVA 1.5 enhances multimodal AI by integrating CLIP-ViT-L-336px with MLP projection and academic VQA data, achieving state-of-the-art results.
0Oct 1, 2023MultimodalBenchmarks
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
LLaVA 1.5 enhances multimodal AI by integrating CLIP-ViT-L-336px with MLP projection and academic VQA data, achieving state-of-the-art results.
Iryna Hartsock, Ghulam Rasool
This review surveys recent medical vision-language models, covering 18 datasets and 16 models for report generation and VQA, highlighting challenges and future directions.