MiniMax

MiniMax M3 (batch)

by MiniMaxtext+image+video->text1 endpoint

Overview

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

Capabilities

Text generation
Image understanding
Video understanding
Tool use / Function calling
Structured output
Extended reasoning
Prompt caching
Long context
Top-P sampling
Stop sequences

Modalities

Input
TextImageVideo
Output
Text

Technical Specifications

Context Window
524.3K tokens
Max Output
471.9K tokens
Tokenizer
Other

Supported Parameters

frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_p

Pricing

Per 1M tokens
Input$0.30
Output$1.20
Cache read$0.06
Added
May 31, 2026
Last synced 9/18/2026