OmniLLM-12B
FreeOpen-source multimodal LLM with trustworthy behavior and strong performance.
FreeFree tier
Inputs: image, textOutputs: text
About OmniLLM-12B
OmniLMM-12B is a multimodal large language model (LMM) developed by OpenBMB, built on EVA02-5B and Zephyr-7B-β with a perceiver resampler layer. It is trained on multimodal data using a curriculum learning approach. The model excels in visual question answering and offers strong performance across multiple benchmarks (MME, MMBench, SEED-Bench). It is designed for trustworthy behavior, being the first open-source LMM to align via multimodal RLHF (RLHF-V), achieving top rankings on MMHal-Bench and outperforming GPT-4V on Object HalBench. Additionally, it supports real-time multimodal interaction.
Key Features
Built on EVA02-5B vision encoder and Zephyr-7B-β language model
Perceiver resampler connector for multimodal fusion
Curriculum learning training on multimodal data
State-of-the-art performance on MME, MMBench, SEED-Bench benchmarks
Multimodal RLHF (RLHF-V) alignment for trustworthy generation
Ranks #1 among open-source models on MMHal-Bench
Outperforms GPT-4V on Object HalBench
Real-time multimodal interaction capability
Pros & Cons
Pros
- Open-source and free to use
- Competitive performance compared to models of similar size
- First open-source LMM aligned with multimodal RLHF for trustworthiness
- Reduces hallucination in image-grounded text generation
- Supports real-time interaction
Cons
- Requires computational resources for inference and fine-tuning
- Limited documentation and community support compared to larger ecosystems
- Dependent on Hugging Face ecosystem for deployment
Best For
Visual question answeringMultimodal dialogue systemsImage-grounded text generation with reduced hallucinationBenchmarking and research in trustworthy AI
FAQ
What is OmniLMM-12B?
OmniLMM-12B is an open-source multimodal large language model (LMM) developed by OpenBMB, based on EVA02-5B and Zephyr-7B-β, designed for visual question answering and real-time multimodal interaction.
How does OmniLMM-12B ensure trustworthy behavior?
It uses a technique called multimodal RLHF (RLHF-V) to align the model, reducing hallucinations and generating text that is factually grounded in images. It ranks #1 among open-source models on MMHal-Bench and outperforms GPT-4V on Object HalBench.
Is OmniLMM-12B free to use?
Yes, the model is open-source and available for free on Hugging Face under the OpenBMB organization.