All Comparisons

GPT-4.1 vs GPT-4o

GPT-4.1
OpenAI 🇺🇸
32.8
#238
high
$2.00/M in · $8.00/M out
1.0M context
GPT-4o
OpenAI 🇺🇸
58.3
#136
high
$2.50/M in · $10.00/M out
128K context
Most Affordable
GPT-4.1
$2.00/M input
Largest Context
GPT-4.1
1.0M tokens
Best Benchmark Score
GPT-4o
58.3/100

Model Specifications

SpecGPT-4.1GPT-4o
ProviderOpenAIOpenAI
Release DateApr 2025May 2024
Knowledge Cutoff2024-06-01Oct 2023
ParametersUndisclosedunknown
Context Window1.0M128K
Max Output
Open SourceNoNo
Licenseproprietaryproprietary
Tokenizer
Modality
Reasoning ModelNoNo
ModeratedNoNo

Pricing Comparison

Price (per 1M tokens)GPT-4.1GPT-4o
Input$2.00$2.50
Output$8.00$10.00

API Provider Pricing

Capabilities

FeatureGPT-4.1GPT-4o
Text Input
Image Input (Vision)
Audio Input
Video Input
File Input
Image Output
Audio Output
Tool Use
Structured Output (JSON)
Streaming
Reasoning Tokens
Web Search
Temperature Control
Top-P Sampling
Stop Sequences
Seed (Reproducibility)
Log Probabilities

Category Comparison

Benchmark-by-Benchmark

general

BenchmarkGPT-4.1GPT-4o
MMMLU
87.3
MMLU-Pro
74.5
IFEval
83.6
Chatbot Arena Elo
1285

coding

BenchmarkGPT-4.1GPT-4o
SWE-bench Verified
54.6
33.2
HumanEval
90.2

math

BenchmarkGPT-4.1GPT-4o
AIME 2025
46.4
AIME 2024
13.4
MATH
76.6
GSM8K
95.8

reasoning

BenchmarkGPT-4.1GPT-4o
GPQA Diamond
66.3
53.6
Humanity's Last Exam
5.4
BigBench-Hard
87.3
ARC-Challenge
96.4

multimodal

BenchmarkGPT-4.1GPT-4o
MMMU
74.8
69.1
CharXiv Reasoning
56.7

agent

BenchmarkGPT-4.1GPT-4o
TAU-Bench Retail
68

language

BenchmarkGPT-4.1GPT-4o
HellaSwag
96.4
WinoGrande
85.7

safety

BenchmarkGPT-4.1GPT-4o
TruthfulQA
73.5

Category Winners

coding
GPT-4o
38.2
math
GPT-4o
54.5
reasoning
GPT-4o
54.4
general
GPT-4o
70.5
language
GPT-4o
79.9
multimodal
GPT-4o
54.0
safety
GPT-4o
94.3
agent
GPT-4.1
50.7

Compare More Models

Frequently Asked Questions