Groq logo

Groq

Free

A cloud inference API for running open-source LLMs, powered by custom LPU hardware.

FreeFree tier
Inputs: text, audioOutputs: text, audio
Type
Open Source
Company
Groq

About Groq

Groq provides a cloud inference API for running open-source large language models (LLMs) and other AI models, powered by its custom Language Processing Unit (LPU) hardware. The LPU is purpose-built for inference, delivering exceptional speed and affordability at scale. GroqCloud offers developers an OpenAI-compatible API with just two lines of code, enabling fast token generation (e.g., 840 tokens per second for Llama 3.1 8B) at low cost (e.g., $0.05 per million input tokens). The platform supports text, audio (TTS, ASR), and built-in tools like web search and code execution. Pricing is linear and predictable with no hidden costs, and a free tier is available. Groq is trusted by enterprises like the McLaren F1 Team for real-time, low-latency AI inference globally.

Key Features

Custom LPU silicon purpose-built for inference since 2016
High-speed token generation, e.g., 840 TPS for Llama 3.1 8B
Low cost per token, e.g., $0.05 per million input tokens
OpenAI-compatible API with two-line code integration
GroqCloud developer platform with scalable infrastructure
Prompt caching for discounted cached input tokens
Built-in tools: web search, code execution, and more
Supports text, text-to-speech, and automatic speech recognition models

Pros & Cons

Pros
  • Fast inference speeds due to custom LPU hardware
  • Cost-effective and predictable pricing without surprise bills
  • Free tier available for developers to get started
  • OpenAI-compatible API simplifies integration
  • Trusted by major clients like McLaren F1 Team
Cons
  • Limited to models offered on the Groq platform (not all open-source models available)
  • May require migration from GPU-based infrastructure
  • Pricing varies by model and may be complex for some use cases

Best For

Real-time AI applications and chatbotsEnterprise inference at scale with predictable costsText-to-speech and speech recognition servicesLatency-sensitive AI deploymentsAI-powered customer support and conversational agents