PreprintarXiv.org2024
Smol VLM
Andrés Marafioti, Orr Zohar, Miquel Farr'e, et al.
SmolVLM introduces a family of compact vision-language models achieving strong performance on resource-constrained devices through aggressive token compression and optimized encoder-LM balance.
250Nov 1, 2024Computer Vision
arXiv