All Comparisons

GPT-4.1 vs Llama 4 Maverick

GPT-4.1
OpenAI πŸ‡ΊπŸ‡Έ
32.8
#238β€”
high
$2.00/M in Β· $8.00/M out
1.0M context
M
Llama 4 Maverick
Meta πŸ‡ΊπŸ‡Έ
52.1
#160β€”
low
Most Affordable
GPT-4.1
$2.00/M input
Largest Context
GPT-4.1
1.0M tokens
Best Benchmark Score
Llama 4 Maverick
52.1/100

Model Specifications

SpecGPT-4.1Llama 4 Maverick
ProviderOpenAIMeta
Release DateApr 2025Apr 2025
Knowledge Cutoff2024-06-01β€”
ParametersUndisclosed400000000000B
Context Window1.0Mβ€”
Max Outputβ€”β€”
Open SourceNoYes
Licenseproprietaryllama_4_community_license_agreement
Tokenizerβ€”β€”
Modalityβ€”β€”
Reasoning ModelNoNo
ModeratedNoNo

Pricing Comparison

Price (per 1M tokens)GPT-4.1Llama 4 Maverick
Input$2.00β€”
Output$8.00β€”

API Provider Pricing

Capabilities

FeatureGPT-4.1Llama 4 Maverick
Text Input
Image Input (Vision)
Audio Input
Video Input
File Input
Image Output
Audio Output
Tool Use
Structured Output (JSON)
Streaming
Reasoning Tokens
Web Search
Temperature Control
Top-P Sampling
Stop Sequences
Seed (Reproducibility)
Log Probabilities

Category Comparison

Benchmark-by-Benchmark

general

BenchmarkGPT-4.1Llama 4 Maverick
MMMLU
87.3
β€”

coding

BenchmarkGPT-4.1Llama 4 Maverick
SWE-bench Verified
54.6
β€”

math

BenchmarkGPT-4.1Llama 4 Maverick
AIME 2025
46.4
β€”

reasoning

BenchmarkGPT-4.1Llama 4 Maverick
GPQA Diamond
66.3
69.8
Humanity's Last Exam
5.4
β€”

multimodal

BenchmarkGPT-4.1Llama 4 Maverick
MMMU
74.8
73.4
CharXiv Reasoning
56.7
β€”
MMMU-Proβ€”
59.6

agent

BenchmarkGPT-4.1Llama 4 Maverick
TAU-Bench Retail
68
β€”

Category Winners

coding
GPT-4.1
25.1
math
GPT-4.1
6.0
reasoning
M
Llama 4 Maverick
54.8
general
GPT-4.1
66.0
multimodal
M
Llama 4 Maverick
46.7
agent
GPT-4.1
50.7

Compare More Models

Frequently Asked Questions