All Comparisons
Claude Opus 4 vs Grok-3
Claude Opus 4
Anthropic 🇺🇸
74.8
#59—
high$15.00/M in · $75.00/M out
200K context
X
Grok-3
xAI 🇺🇸
77.5
#44—
low$3.00/M in · $15.00/M out
128K context
Most Affordable
Grok-3
$3.00/M input
Largest Context
Claude Opus 4
200K tokens
Best Benchmark Score
Grok-3
77.5/100
Model Specifications
| Spec | Claude Opus 4 | Grok-3 |
|---|---|---|
| Provider | Anthropic | xAI |
| Release Date | May 2025 | Feb 2025 |
| Knowledge Cutoff | Mar 2025 | 2024-11-17 |
| Parameters | unknown | Undisclosed |
| Context Window | 200K | 128K |
| Max Output | — | — |
| Open Source | No | No |
| License | proprietary | proprietary |
| Tokenizer | — | — |
| Modality | — | — |
| Reasoning Model | No | No |
| Moderated | No | No |
Pricing Comparison
| Price (per 1M tokens) | Claude Opus 4 | Grok-3 |
|---|---|---|
| Input | $15.00 | $3.00 |
| Output | $75.00 | $15.00 |
API Provider Pricing
Capabilities
| Feature | Claude Opus 4 | Grok-3 |
|---|---|---|
| Text Input | ||
| Image Input (Vision) | ||
| Audio Input | ||
| Video Input | ||
| File Input | ||
| Image Output | ||
| Audio Output | ||
| Tool Use | ||
| Structured Output (JSON) | ||
| Streaming | ||
| Reasoning Tokens | ||
| Web Search | ||
| Temperature Control | ||
| Top-P Sampling | ||
| Stop Sequences | ||
| Seed (Reproducibility) | ||
| Log Probabilities |
Category Comparison
Benchmark-by-Benchmark
general
| Benchmark | Claude Opus 4 | Grok-3 |
|---|---|---|
| Chatbot Arena Elo | 1365 | — |
| IFEval | 91.5 | — |
| MMLU-Pro | 83.5 | — |
coding
| Benchmark | Claude Opus 4 | Grok-3 |
|---|---|---|
| SWE-bench Verified | 55.8 | — |
| HumanEval | 95.1 | — |
reasoning
| Benchmark | Claude Opus 4 | Grok-3 |
|---|---|---|
| GPQA Diamond | 74.8 | 84.6 |
| BigBench-Hard | 93.2 | — |
multimodal
| Benchmark | Claude Opus 4 | Grok-3 |
|---|---|---|
| MMMU | 74.2 | 78 |
Category Winners
coding
58.0
math
75.5
reasoning
X
Grok-380.0
general
96.0
multimodal
X
Grok-375.9