All Comparisons
Grok-3 vs o3
X
Grok-3
xAI πΊπΈ
77.9
#38β
low$3.00/M in Β· $15.00/M out
128K context
o3
OpenAI πΊπΈ
47.9
#174β
highMost Affordable
Grok-3
$3.00/M input
Largest Context
Grok-3
128K tokens
Best Benchmark Score
Grok-3
77.9/100
Model Specifications
| Spec | Grok-3 | o3 |
|---|---|---|
| Provider | xAI | OpenAI |
| Release Date | Feb 2025 | Apr 2025 |
| Knowledge Cutoff | 2024-11-17 | 2024-05-31 |
| Parameters | Undisclosed | Undisclosed |
| Context Window | 128K | β |
| Max Output | β | β |
| Open Source | No | No |
| License | proprietary | proprietary |
| Tokenizer | β | β |
| Modality | β | β |
| Reasoning Model | No | No |
| Moderated | No | No |
Pricing Comparison
| Price (per 1M tokens) | Grok-3 | o3 |
|---|---|---|
| Input | $3.00 | β |
| Output | $15.00 | β |
API Provider Pricing
Capabilities
| Feature | Grok-3 | o3 |
|---|---|---|
| Text Input | ||
| Image Input (Vision) | ||
| Audio Input | ||
| Video Input | ||
| File Input | ||
| Image Output | ||
| Audio Output | ||
| Tool Use | ||
| Structured Output (JSON) | ||
| Streaming | ||
| Reasoning Tokens | ||
| Web Search | ||
| Temperature Control | ||
| Top-P Sampling | ||
| Stop Sequences | ||
| Seed (Reproducibility) | ||
| Log Probabilities |
Category Comparison
Benchmark-by-Benchmark
math
| Benchmark | Grok-3 | o3 |
|---|---|---|
| AIME 2025 | 93.3 | 86.4 |
| FrontierMath | β | 15.8 |
reasoning
| Benchmark | Grok-3 | o3 |
|---|---|---|
| GPQA Diamond | 84.6 | 83.3 |
| Humanity's Last Exam | β | 14.7 |
| ARC-AGI v2 | β | 6.5 |
multimodal
| Benchmark | Grok-3 | o3 |
|---|---|---|
| MMMU | 78 | 82.9 |
| MMMU-Pro | β | 76.4 |
| CharXiv Reasoning | β | 78.6 |
coding
| Benchmark | Grok-3 | o3 |
|---|---|---|
| SWE-bench Verified | β | 69.1 |
agent
| Benchmark | Grok-3 | o3 |
|---|---|---|
| BrowseComp | β | 49.7 |
Category Winners
coding
58.4
math
X
Grok-375.2
reasoning
X
Grok-380.6
multimodal
X
Grok-376.7
agent
23.4