All Comparisons

DeepSeek-V3 vs GPT-4o
DeepSeek-V3
DeepSeek 🇨🇳
23.4
#265—
lowGPT-4o
OpenAI 🇺🇸
58.3
#136—
high$2.50/M in · $10.00/M out
128K context
Most Affordable
GPT-4o
$2.50/M input
Largest Context
GPT-4o
128K tokens
Best Benchmark Score
GPT-4o
58.3/100
Model Specifications
| Spec | DeepSeek-V3 | GPT-4o |
|---|---|---|
| Provider | DeepSeek | OpenAI |
| Release Date | Dec 2024 | May 2024 |
| Knowledge Cutoff | — | Oct 2023 |
| Parameters | 671000000000B | unknown |
| Context Window | — | 128K |
| Max Output | — | — |
| Open Source | Yes | No |
| License | mit_+_model_license_(commercial_use_allowed) | proprietary |
| Tokenizer | — | — |
| Modality | — | — |
| Reasoning Model | No | No |
| Moderated | No | No |
Pricing Comparison
| Price (per 1M tokens) | DeepSeek-V3 | GPT-4o |
|---|---|---|
| Input | — | $2.50 |
| Output | — | $10.00 |
API Provider Pricing
Capabilities
| Feature | DeepSeek-V3 | GPT-4o |
|---|---|---|
| Text Input | ||
| Image Input (Vision) | ||
| Audio Input | ||
| Video Input | ||
| File Input | ||
| Image Output | ||
| Audio Output | ||
| Tool Use | ||
| Structured Output (JSON) | ||
| Streaming | ||
| Reasoning Tokens | ||
| Web Search | ||
| Temperature Control | ||
| Top-P Sampling | ||
| Stop Sequences | ||
| Seed (Reproducibility) | ||
| Log Probabilities |
Category Comparison
Benchmark-by-Benchmark
general
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| SimpleQA | 24.9 | — |
| MMLU-Pro | — | 74.5 |
| IFEval | — | 83.6 |
| Chatbot Arena Elo | — | 1285 |
coding
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| SWE-bench Verified | 42 | 33.2 |
| HumanEval | — | 90.2 |
reasoning
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| GPQA Diamond | 59.1 | 53.6 |
| BigBench-Hard | — | 87.3 |
| ARC-Challenge | — | 96.4 |
language
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| HellaSwag | — | 96.4 |
| WinoGrande | — | 85.7 |
multimodal
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| MMMU | — | 69.1 |
safety
| Benchmark | DeepSeek-V3 | GPT-4o |
|---|---|---|
| TruthfulQA | — | 73.5 |
Category Winners
coding
38.2
math
54.5
reasoning
54.4
general
70.5
language
79.9
multimodal
54.0
safety
94.3