Preprint2023
PaLI-3
Unknown
A 5B vision-language model using a 2B SigLIP vision encoder and 3B UL2 language model achieves SOTA on video QA without any video pretraining.
0Oct 1, 2023BenchmarksComputer Vision
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.