Gemma2-9|27B logo

Gemma2-9|27B

Free

Next-gen open language model with best-in-class performance and efficiency

FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
Google

About Gemma2-9|27B

Gemma 2 is Google's next-generation family of open language models, launched in June 2024. Available in 9 billion and 27 billion parameter sizes, Gemma 2 is built from the same research and technology as the Gemini models. It offers best-in-class performance for its size class, with the 27B model delivering competitive alternatives to models more than twice its size. Designed for efficiency, it runs inference at full precision on a single Google Cloud TPU host or NVIDIA A100/H100 GPU, significantly reducing deployment costs. The model is optimized for fast inference across diverse hardware, from gaming laptops to cloud servers, and integrates with major AI frameworks like Hugging Face Transformers, JAX, PyTorch, TensorFlow, vLLM, Gemma.cpp, Llama.cpp, and Ollama. Gemma 2 is released under the commercially-friendly Gemma license, enabling sharing and commercialization, and includes significant safety advancements. It can be fine-tuned with Keras and Hugging Face, and is optimized with NVIDIA TensorRT-LLM for accelerated infrastructure.

Key Features

Outsized performance: 27B model competitive with models twice its size; 9B outperforms Llama 3 8B
Unmatched efficiency: Runs inference at full precision on a single Google Cloud TPU host or NVIDIA A100/H100 GPU
Blazing fast inference across hardware: optimized for gaming laptops, desktops, and cloud setups
Open and accessible: commercially-friendly Gemma license allows sharing and commercialization
Broad framework compatibility: supports Hugging Face Transformers, JAX, PyTorch, TensorFlow, vLLM, Gemma.cpp, Llama.cpp, Ollama, NVIDIA TensorRT-LLM
Effortless deployment: Google Cloud deployment upcoming, easy fine-tuning with Keras and Hugging Face
Built-in safety advancements from Google DeepMind research

Pros & Cons

Pros
  • Competitive performance for its size, rivaling larger proprietary models
  • Efficient inference on single GPU/TPU, reducing deployment costs
  • Commercially-friendly open license encourages innovation and sharing
  • Broad compatibility with popular AI frameworks and tools
  • Built with significant safety advancements from Google DeepMind
Cons
  • Limited to 9B and 27B parameter sizes; may not match capabilities of larger frontier models
  • Optimal performance requires NVIDIA or Google Cloud hardware for full precision

Best For

Building and deploying AI applications with open language modelsRunning efficient inference on local hardware (gaming laptops, desktops) or cloud infrastructureFine-tuning for domain-specific tasks using Keras and Hugging FaceResearch and experimentation with state-of-the-art open modelsPrototyping AI products that require high performance at low cost

FAQ

What is Gemma 2?
Gemma 2 is Google's next-generation family of open language models, available in 9 billion and 27 billion parameter sizes, built from the research used for Gemini models.
What sizes of Gemma 2 are available?
Gemma 2 is available in 9B and 27B parameter sizes.
How can I use Gemma 2?
Gemma 2 can be used via Google AI Studio, quantized version with Gemma.cpp on CPU, or on home computers with NVIDIA RTX GPUs via Hugging Face Transformers. It integrates with many frameworks including JAX, PyTorch, TensorFlow, vLLM, Ollama, and more.
What is the license for Gemma 2?
Gemma 2 is released under the commercially-friendly Gemma license, allowing sharing and commercialization.
Can Gemma 2 be fine-tuned?
Yes, Gemma 2 can be fine-tuned with Keras and Hugging Face, with additional parameter-efficient fine-tuning options being developed.