All Comparisons

Command R+ vs GPT-4.1
Command R+
Cohere 🇨🇦
8.5
#316—
high$2.50/M in · $10.00/M out
128K context
GPT-4.1
OpenAI 🇺🇸
32.8
#238—
high$2.00/M in · $8.00/M out
1.0M context
Most Affordable
GPT-4.1
$2.00/M input
Largest Context
GPT-4.1
1.0M tokens
Best Benchmark Score
GPT-4.1
32.8/100
Model Specifications
| Spec | Command R+ | GPT-4.1 |
|---|---|---|
| Provider | Cohere | OpenAI |
| Release Date | Apr 2024 | Apr 2025 |
| Knowledge Cutoff | — | 2024-06-01 |
| Parameters | 104B | Undisclosed |
| Context Window | 128K | 1.0M |
| Max Output | — | — |
| Open Source | Open Weights | No |
| License | cc-by-nc-4.0 | proprietary |
| Tokenizer | — | — |
| Modality | — | — |
| Reasoning Model | No | No |
| Moderated | No | No |
Pricing Comparison
| Price (per 1M tokens) | Command R+ | GPT-4.1 |
|---|---|---|
| Input | $2.50 | $2.00 |
| Output | $10.00 | $8.00 |
API Provider Pricing
Capabilities
| Feature | Command R+ | GPT-4.1 |
|---|---|---|
| Text Input | ||
| Image Input (Vision) | ||
| Audio Input | ||
| Video Input | ||
| File Input | ||
| Image Output | ||
| Audio Output | ||
| Tool Use | ||
| Structured Output (JSON) | ||
| Streaming | ||
| Reasoning Tokens | ||
| Web Search | ||
| Temperature Control | ||
| Top-P Sampling | ||
| Stop Sequences | ||
| Seed (Reproducibility) | ||
| Log Probabilities |
Category Comparison
Benchmark-by-Benchmark
general
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| IFEval | 73.8 | — |
| Chatbot Arena Elo | 1168 | — |
| MMLU-Pro | 55.8 | — |
| MMMLU | — | 87.3 |
coding
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| HumanEval | 73.2 | — |
| SWE-bench Verified | — | 54.6 |
reasoning
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| ARC-Challenge | 88.4 | — |
| GPQA Diamond | — | 66.3 |
| Humanity's Last Exam | — | 5.4 |
language
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| WinoGrande | 80.2 | — |
| HellaSwag | 85.3 | — |
safety
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| TruthfulQA | 58.8 | — |
multimodal
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| MMMU | — | 74.8 |
| CharXiv Reasoning | — | 56.7 |
agent
| Benchmark | Command R+ | GPT-4.1 |
|---|---|---|
| TAU-Bench Retail | — | 68 |
Category Winners
coding
25.1
math
6.0
reasoning
27.5
general
66.0
language
6.8
multimodal
40.2
safety
21.0
agent
50.7