All Comparisons


GPT-4o vs GPT-4.1 vs o3
GPT-4o
OpenAI 🇺🇸
58.1
#145—
high$2.50/M in · $10.00/M out
128K context
GPT-4.1
OpenAI 🇺🇸
32.5
#256—
high$2.00/M in · $8.00/M out
1.0M context
o3
OpenAI 🇺🇸
47.8
#187—
high$2.00/M in · $8.00/M out
200K context
Most Affordable
GPT-4.1
$2.00/M input
Largest Context
GPT-4.1
1.0M tokens
Best Benchmark Score
GPT-4o
58.1/100
Model Specifications
| Spec | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| Provider | OpenAI | OpenAI | OpenAI |
| Release Date | May 2024 | Apr 2025 | Apr 2025 |
| Knowledge Cutoff | 2023-10-31 | 2024-06-30 | 2024-06-30 |
| Parameters | unknown | Undisclosed | Undisclosed |
| Context Window | 128K | 1.0M | 200K |
| Max Output | 16.4K | 32.8K | 100K |
| Open Source | No | No | No |
| License | proprietary | proprietary | proprietary |
| Tokenizer | GPT | GPT | GPT |
| Modality | text+image+file->text | text+image+file->text | text+image+file->text |
| Reasoning Model | No | No | No |
| Moderated | Yes | Yes | Yes |
Pricing Comparison
| Price (per 1M tokens) | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| Input | $2.50 | $2.00 | $2.00 |
| Output | $10.00 | $8.00 | $8.00 |
| Cache Read | $1.25 | $0.50 | $0.50 |
API Provider Pricing
Capabilities
| Feature | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| Text Input | |||
| Image Input (Vision) | |||
| Audio Input | |||
| Video Input | |||
| File Input | |||
| Image Output | |||
| Audio Output | |||
| Tool Use | |||
| Structured Output (JSON) | |||
| Streaming | |||
| Reasoning Tokens | |||
| Web Search | |||
| Temperature Control | |||
| Top-P Sampling | |||
| Stop Sequences | |||
| Seed (Reproducibility) | |||
| Log Probabilities |
Category Comparison
Benchmark-by-Benchmark
general
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| MMLU-Pro | 74.5 | — | — |
| Chatbot Arena Elo | 1285 | — | — |
| IFEval | 83.6 | — | — |
| MMMLU | — | 87.3 | — |
coding
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| HumanEval | 90.2 | — | — |
| SWE-bench Verified | 33.2 | 54.6 | 69.1 |
math
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| GSM8K | 95.8 | — | — |
| MATH | 76.6 | — | — |
| AIME 2024 | 13.4 | — | — |
| AIME 2025 | — | 46.4 | 86.4 |
| FrontierMath | — | — | 15.8 |
reasoning
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| BigBench-Hard | 87.3 | — | — |
| ARC-Challenge | 96.4 | — | — |
| GPQA Diamond | 53.6 | 66.3 | 83.3 |
| Humanity's Last Exam | — | 5.4 | 14.7 |
| ARC-AGI v2 | — | — | 6.5 |
language
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| WinoGrande | 85.7 | — | — |
| HellaSwag | 96.4 | — | — |
multimodal
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| MMMU | 69.1 | 74.8 | 82.9 |
| CharXiv Reasoning | — | 56.7 | 78.6 |
| MMMU-Pro | — | — | 76.4 |
safety
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| TruthfulQA | 73.5 | — | — |
agent
| Benchmark | GPT-4o | GPT-4.1 | o3 |
|---|---|---|---|
| TAU-Bench Retail | — | 68 | — |
| BrowseComp | — | — | 49.7 |
Category Winners
coding
58.4
math
54.5
reasoning
54.2
general
70.5
language
79.9
multimodal
72.8
safety
94.3
agent
50.7