All Comparisons

GPT-4o vs GPT-4.1 vs o3

GPT-4o
OpenAI 🇺🇸
58.1
#145
high
$2.50/M in · $10.00/M out
128K context
GPT-4.1
OpenAI 🇺🇸
32.5
#256
high
$2.00/M in · $8.00/M out
1.0M context
o3
OpenAI 🇺🇸
47.8
#187
high
$2.00/M in · $8.00/M out
200K context
Most Affordable
GPT-4.1
$2.00/M input
Largest Context
GPT-4.1
1.0M tokens
Best Benchmark Score
GPT-4o
58.1/100

Model Specifications

SpecGPT-4oGPT-4.1o3
ProviderOpenAIOpenAIOpenAI
Release DateMay 2024Apr 2025Apr 2025
Knowledge Cutoff2023-10-312024-06-302024-06-30
ParametersunknownUndisclosedUndisclosed
Context Window128K1.0M200K
Max Output16.4K32.8K100K
Open SourceNoNoNo
Licenseproprietaryproprietaryproprietary
TokenizerGPTGPTGPT
Modalitytext+image+file->texttext+image+file->texttext+image+file->text
Reasoning ModelNoNoNo
ModeratedYesYesYes

Pricing Comparison

Price (per 1M tokens)GPT-4oGPT-4.1o3
Input$2.50$2.00$2.00
Output$10.00$8.00$8.00
Cache Read$1.25$0.50$0.50

API Provider Pricing

Capabilities

FeatureGPT-4oGPT-4.1o3
Text Input
Image Input (Vision)
Audio Input
Video Input
File Input
Image Output
Audio Output
Tool Use
Structured Output (JSON)
Streaming
Reasoning Tokens
Web Search
Temperature Control
Top-P Sampling
Stop Sequences
Seed (Reproducibility)
Log Probabilities

Category Comparison

Benchmark-by-Benchmark

general

BenchmarkGPT-4oGPT-4.1o3
MMLU-Pro
74.5
Chatbot Arena Elo
1285
IFEval
83.6
MMMLU
87.3

coding

BenchmarkGPT-4oGPT-4.1o3
HumanEval
90.2
SWE-bench Verified
33.2
54.6
69.1

math

BenchmarkGPT-4oGPT-4.1o3
GSM8K
95.8
MATH
76.6
AIME 2024
13.4
AIME 2025
46.4
86.4
FrontierMath
15.8

reasoning

BenchmarkGPT-4oGPT-4.1o3
BigBench-Hard
87.3
ARC-Challenge
96.4
GPQA Diamond
53.6
66.3
83.3
Humanity's Last Exam
5.4
14.7
ARC-AGI v2
6.5

language

BenchmarkGPT-4oGPT-4.1o3
WinoGrande
85.7
HellaSwag
96.4

multimodal

BenchmarkGPT-4oGPT-4.1o3
MMMU
69.1
74.8
82.9
CharXiv Reasoning
56.7
78.6
MMMU-Pro
76.4

safety

BenchmarkGPT-4oGPT-4.1o3
TruthfulQA
73.5

agent

BenchmarkGPT-4oGPT-4.1o3
TAU-Bench Retail
68
BrowseComp
49.7

Category Winners

coding
o3
58.4
math
GPT-4o
54.5
reasoning
GPT-4o
54.2
general
GPT-4o
70.5
language
GPT-4o
79.9
multimodal
o3
72.8
safety
GPT-4o
94.3
agent
GPT-4.1
50.7

Frequently Asked Questions