Model Providers: Configuration and CLI Reference
Learn how to configure LLM providers, set model refs, and manage auth via CLI. Essential for developers setting up model access and defaults.
Read this when
- You need a provider-by-provider model setup reference
- You want example configs or CLI onboarding commands for model providers
Reference for LLM/model providers (not chat channels like WhatsApp/Telegram). For model selection rules, see Models.
Quick rules
Model refs and CLI helpers
- Model refs follow the
provider/modelformat (for instance,opencode/claude-opus-4-6). - Aliases and model-specific settings are held in
agents.defaults.models; the optional explicit override allowlist isagents.defaults.modelPolicy.allow. - Command-line helpers:
openclaw onboard,openclaw models list,openclaw models set <provider/model>. - The provider-level output-token default is set by
models.providers.*.maxTokens. Within eachmodels.providers.*.models[]entry,contextWindowspecifies the native window,contextTokenslimits active input, andmaxTokensoverrides that model's output capacity. - For fallback rules, cooldown probes, and session-override persistence, refer to Model failover.
Adding provider auth does not change your primary model
When you add or reauth a provider, openclaw configure keeps an existing agents.defaults.model.primary intact. openclaw models auth login behaves the same way unless --set-default is supplied. Provider plugins may still suggest a default model in their auth config patch, but OpenClaw interprets that as "make this model available" when a primary model is already present, not "replace the current primary model."
To deliberately change the default model, run openclaw models set <provider/model> or openclaw models auth login --provider <id> --set-default.
OpenAI provider/runtime split
OpenAI model refs and agent runtimes are distinct concepts:
- The canonical OpenAI provider and model are chosen via
openai/<model>. A prefix alone never selects Codex. - When provider/model runtime policy is unset or set to
auto, OpenAI may implicitly pick Codex only for an exact official HTTPS Platform Responses or ChatGPT Responses route with no authored provider request override. Valid model-scoped Fast-mode controls do not qualify as authored request params. - Authored Completions adapters, custom endpoints, and routes with authored request behavior remain on OpenClaw. Plaintext official HTTP endpoints are rejected.
- Legacy Codex model refs are legacy config that doctor rewrites to
openai/<model>. - An otherwise eligible route is explicitly kept on OpenClaw by provider/model
agentRuntime.id: "openclaw".agentRuntime.id: "codex"demands Codex and fails closed when the effective route is not Codex-compatible.
Check OpenAI implicit agent runtime and Codex harness. If the provider/runtime split is unclear, start with Agent runtimes.
Plugin auto-enable follows the same boundary: an implicitly Codex-compatible effective route can enable the Codex plugin, while explicit provider/model agentRuntime.id: "codex" or legacy codex/<model> refs require it. An openai/* prefix alone does not.
Fresh OpenAI API-key and ChatGPT/Codex OAuth setup select the canonical
openai/gpt-5.6-sol ref. The bare direct-API openai/gpt-5.6 alias remains
supported and resolves to Sol. Existing explicit primaries, including
openai/gpt-5.5, are preserved when OpenAI auth is added or refreshed. GPT-5.5 remains available
through either runtime as an explicit recovery choice for accounts without
GPT-5.6 access.
CLI runtimes
CLI runtimes use the same split: choose canonical model refs such as anthropic/claude-* or google/gemini-*, then set provider/model runtime policy to claude-cli or google-gemini-cli when you want a local CLI backend.
Legacy claude-cli/* and google-gemini-cli/* refs migrate back to canonical provider refs with the runtime recorded separately. Legacy codex-cli/* refs migrate to openai/* and use the Codex app-server route; OpenClaw no longer keeps a bundled Codex CLI backend.
Configure providers in the Control UI
Open Settings → Model Providers in the Control UI to add, replace, or remove provider API keys stored in models.providers.<id>.apiKey. The page identifies whether each API key comes from OpenClaw config or an environment variable without displaying the credential. Environment-provided keys remain managed by the gateway process environment.
Use Test connection to run a live provider probe and see latency or a categorized authentication, rate-limit, billing, timeout, or response error. A probe makes a real provider request and may consume a small number of tokens. OAuth and token profiles can also be logged out from the provider card.
The Default models card manages the primary model, ordered fallbacks, and utility model from the configured model catalog. Choose the models, then save them together to the existing agents.defaults.model and agents.defaults.utilityModel settings. For the utility model, Automatic leaves the setting unset and Disabled stores an empty string to turn utility routing off.
Plugin-owned provider behavior
Most provider-specific logic lives in provider plugins (registerProvider(...)) while OpenClaw keeps the generic inference loop. Plugins own onboarding, model catalogs, auth env-var mapping, transport/config normalization, tool-schema cleanup, failover classification, OAuth refresh, usage reporting, thinking/reasoning profiles, and more.
The full list of provider-SDK hooks and bundled-plugin examples lives in Provider plugins. A provider that needs a totally custom request executor is a separate, deeper extension surface.
Note
Provider-owned runner behavior lives on explicit provider hooks such as replay policy, tool-schema normalization, stream wrapping, and transport/request helpers. The legacy
ProviderPlugin.capabilitiesstatic bag is compatibility-only and is no longer read by shared runner logic.
API key rotation
Key sources and priority
Configure multiple keys via:
OPENCLAW_LIVE_<PROVIDER>_KEY(single live override, highest priority)<PROVIDER>_API_KEYS(comma or semicolon list)<PROVIDER>_API_KEY(primary key)<PROVIDER>_API_KEY_*(numbered list, e.g.<PROVIDER>_API_KEY_1)
For Google providers, GOOGLE_API_KEY is additionally included as a fallback. The key selection order maintains priority and removes duplicate values.
When rotation kicks in
- Retries with the next key happen only on rate-limit responses (such as
429,rate_limit,quota,resource exhausted,Too many concurrent requests,ThrottlingException,concurrency limit reached,workers_ai ... quota limit exceeded, or recurring usage-limit notices). - Failures that are not rate-limit related stop immediately; no key rotation is attempted.
- If every candidate key fails, the error from the last attempt is returned as the final result.
Official provider plugins
Official provider plugins publish their own model catalog rows. These providers do not require any models.providers model entries; enable the provider plugin, configure auth, and select a model. Use models.providers only for explicit custom providers or narrow request settings like timeouts.
OpenAI
- Provider:
openai - Auth:
OPENAI_API_KEY - Optional rotation:
OPENAI_API_KEYS,OPENAI_API_KEY_1,OPENAI_API_KEY_2, plusOPENCLAW_LIVE_OPENAI_KEY(single override) - Fresh setup default:
openai/gpt-5.6-sol. - Example models:
openai/gpt-5.6-sol,openai/gpt-5.6-terra,openai/gpt-5.6-luna,openai/gpt-5.5; the bare direct-APIopenai/gpt-5.6alias remains supported. - Verify account/model availability with
openclaw models list --provider openaiif a specific install or API key behaves differently. - CLI:
openclaw onboard --auth-choice openai-api-key - Direct OpenAI API-key Responses requests default to
"sse". - Override per model via
agents.defaults.models["openai/<model>"].params.transport("sse","websocket","websocket-cached", or"auto"). Cached WebSockets reuse the session connection and send only new input withprevious_response_idwhen history still matches. - Set an explicit OpenAI API service tier with
params.serviceTierorparams.service_tier; Fast mode (formerly Priority processing) usesservice_tier=priority. - On native public OpenAI and ChatGPT/Codex Responses requests, precedence is payload/transport
service_tier, then a valid explicit model param, then the fast-mode default. /fastand validparams.fastMode/params.fast_modevalues are shared agent-runtime controls; on direct embeddedopenai/*Responses requests they supplyservice_tier=priorityonly when no higher-precedence tier exists.- Hidden OpenClaw attribution headers (
originator,version,User-Agent) apply only on native OpenAI traffic toapi.openai.com, not generic OpenAI-compatible proxies - Native OpenAI routes also keep Responses
store, prompt-cache hints, and OpenAI reasoning-compat payload shaping; proxy routes do not openai/gpt-5.3-codex-sparkis available only through ChatGPT/Codex OAuth; direct OpenAI API-key and Azure API-key routes reject it
{
agents: { defaults: { model: { primary: "openai/gpt-5.6-sol" } } },
}
If the API organization does not expose GPT-5.6, set
openai/gpt-5.5 explicitly. Normal onboarding and reauthentication preserve an
existing explicit primary model; models auth login --set-default and
models set are the intentional replacement paths.
Anthropic
- Provider:
anthropic - Auth:
ANTHROPIC_API_KEY - Optional rotation:
ANTHROPIC_API_KEYS,ANTHROPIC_API_KEY_1,ANTHROPIC_API_KEY_2, plusOPENCLAW_LIVE_ANTHROPIC_KEY(single override) - Example model:
anthropic/claude-opus-5 - CLI:
openclaw onboard --auth-choice apiKey - Direct public Anthropic requests support the shared
/fasttoggle andparams.fastMode, including API-key and OAuth-authenticated traffic sent toapi.anthropic.com; OpenClaw maps that to Anthropicservice_tier(autovsstandard_only) - Preferred Claude CLI config keeps the model ref canonical and selects the CLI
backend separately:
anthropic/claude-opus-5with model-scopedagentRuntime.id: "claude-cli". Legacyclaude-cli/claude-opus-4-7refs still work for compatibility.
Note
Claude CLI reuse (
claude -p) is a sanctioned OpenClaw integration path. Anthropic setup-token auth remains supported, but OpenClaw prefers Claude CLI reuse when available.
{
agents: { defaults: { model: { primary: "anthropic/claude-opus-5" } } },
}
OpenAI ChatGPT/Codex OAuth
- Provider:
openai - Auth: OAuth (ChatGPT)
- Fresh native Codex app-server harness ref:
openai/gpt-5.6-sol - Native Codex app-server harness docs: Codex harness
- Legacy model refs:
codex/gpt-*,openai-codex/gpt-* - Plugin boundary:
openai/*loads the OpenAI plugin; explicit runtime policy or the provider-owned effective route decides whether the native Codex app-server plugin is selected. - CLI:
openclaw onboard --auth-choice openaioropenclaw models auth login --provider openai - OpenClaw's embedded ChatGPT Responses transport defaults to
auto(WebSocket-first, SSE fallback). agents.defaults.models["openai/<model>"].params.transportandparams.serviceTierare authored embedded-provider request settings. They keep implicit runtime selection on OpenClaw; native Codex owns its app-server transport and service tier.- Valid model-scoped
params.fastMode/params.fast_modevalues and valid cutoff keys are portable typed agent-runtime controls. They do not count as authored provider request params and do not select a runtime. PinagentRuntime.id: "openclaw"oragentRuntime.id: "codex"when a recipe depends on one runtime. - Hidden OpenClaw attribution headers (
originator,version,User-Agent) are only attached on native Codex traffic tochatgpt.com/backend-api, not generic OpenAI-compatible proxies - The shared
/fasttoggle, configured defaults, and valid model-scoped Fast params resolve through one runtime-control policy. See Thinking levels for precedence. - OpenAI API Fast mode is premium-priced and model-specific. GPT-5.6 Sol currently costs 2× Standard token pricing, and long-context multipliers stack. ChatGPT/Codex-credit Fast mode is separate: GPT-5.6 and GPT-5.5 currently consume 2.5× Standard credits, while API-key Codex runs use API token pricing. See Fast mode, API pricing, and Codex speed.
- The native Codex catalog can expose exact
openai/gpt-5.6-sol,openai/gpt-5.6-terra, andopenai/gpt-5.6-lunarefs according to account access. It does not apply the direct API's baregpt-5.6alias client-side. openai/gpt-5.5uses the Codex catalog nativecontextWindow = 400000and default runtimecontextTokens = 272000; override the runtime cap withmodels.providers.openai.models[].contextTokens- Sign in with
openaiauth and useopenai/gpt-5.6-solfor a fresh subscription-backed setup. Selectopenai/gpt-5.5explicitly if that Codex workspace does not expose GPT-5.6. - Use provider/model
agentRuntime.id: "openclaw"to keep an otherwise eligible route on the built-in runtime. With runtime unset orauto, only an exact official HTTPS Responses/ChatGPT-compatible route with no authored provider request override may select Codex implicitly. - Legacy Codex GPT refs are legacy state, not a live provider route. Use canonical
openai/*refs for new agent config, and runopenclaw doctor --fixto migratecodex/*andopenai-codex/*refs while preserving their native Codex semantics with model-scopedagentRuntime.id: "codex". Existing explicit canonicalopenai/gpt-5.5selections are not upgraded.
{
plugins: { entries: { codex: { enabled: true } } },
agents: {
defaults: {
model: { primary: "openai/gpt-5.6-sol" },
},
},
}
{
models: {
providers: {
openai: {
models: [{ id: "gpt-5.5", contextTokens: 160000 }],
},
},
},
}
Other subscription-style hosted options
-
MiniMax, MiniMax Coding Plan OAuth or API key access.
-
Qwen Cloud, Qwen Cloud provider surface plus Alibaba DashScope and Coding Plan endpoint mapping.
-
Z.AI (GLM), Z.AI Coding Plan or general API endpoints.
OpenCode
- Auth:
OPENCODE_API_KEY(orOPENCODE_ZEN_API_KEY) - Zen runtime provider:
opencode - Go runtime provider:
opencode-go - Example models:
opencode/claude-opus-4-6,opencode-go/kimi-k2.6 - CLI:
openclaw onboard --auth-choice opencode-zenoropenclaw onboard --auth-choice opencode-go
{
agents: { defaults: { model: { primary: "opencode/claude-opus-4-6" } } },
}
Google Gemini (API key)
- Provider:
google - Auth:
GEMINI_API_KEY - Optional rotation:
GEMINI_API_KEYS,GEMINI_API_KEY_1,GEMINI_API_KEY_2,GOOGLE_API_KEYfallback, andOPENCLAW_LIVE_GEMINI_KEY(single override) - Example models:
google/gemini-3.1-pro-preview,google/gemini-3.5-flash - Compatibility: legacy OpenClaw config using
google/gemini-3.1-flash-previewis normalized togoogle/gemini-3-flash-preview - Alias:
google/gemini-3.1-prois accepted and normalized to Google's live Gemini API id,google/gemini-3.1-pro-preview - CLI:
openclaw onboard --auth-choice gemini-api-key - Thinking:
/think adaptiveuses Google dynamic thinking. Gemini 3/3.1 omit a fixedthinkingLevel; Gemini 2.5 sendsthinkingBudget: -1. - Direct Gemini runs also accept
agents.defaults.models["google/<model>"].params.cachedContent(or legacycached_content) to forward a provider-nativecachedContents/...handle; Gemini cache hits surface as OpenClawcacheRead
Google Vertex and Gemini CLI runtime
google-vertex: managed Google Cloud access through gcloud Application Default Credentials.google-gemini-cli: optional local runtime for an explicitly configured canonicalgoogle/*model.
OpenClaw does not create Gemini CLI OAuth or Antigravity OAuth profiles. Connect Google through an AI Studio API key or Vertex AI. If you explicitly choose the Gemini CLI runtime, it can use the selected Google API-key profile. Existing valid Gemini CLI OAuth profiles remain runtime-compatible, but they are not a setup or recovery route.
Gemini CLI uses stream-json by default. OpenClaw reads assistant stream
messages and normalizes stats.cached into cacheRead; legacy
--output-format json overrides still read reply text from response.
Z.AI (GLM)
- Provider:
zai - Auth:
ZAI_API_KEY - Example model:
zai/glm-5.2 - CLI:
openclaw onboard --auth-choice zai-api-key- Model refs use the canonical
zai/*provider ID. zai-api-keyauto-detects the matching Z.AI endpoint;zai-coding-global,zai-coding-cn,zai-global, andzai-cnforce a specific surface
- Model refs use the canonical
Vercel AI Gateway
- Provider:
vercel-ai-gateway - Auth:
AI_GATEWAY_API_KEY - Example models:
vercel-ai-gateway/anthropic/claude-opus-4.6,vercel-ai-gateway/moonshotai/kimi-k2.6 - CLI:
openclaw onboard --auth-choice ai-gateway-api-key
Other bundled provider plugins
| Provider | Id | Auth env | Example model |
|---|---|---|---|
| Arcee | arcee | ARCEEAI_API_KEY or OPENROUTER_API_KEY | arcee/trinity-large-thinking |
| BytePlus | byteplus / byteplus-plan | BYTEPLUS_API_KEY | byteplus-plan/ark-code-latest |
| Cerebras | cerebras | CEREBRAS_API_KEY | cerebras/zai-glm-4.7 |
| Chutes | chutes | CHUTES_API_KEY or CHUTES_OAUTH_TOKEN | chutes/zai-org/GLM-5-TEE |
| ClawRouter | clawrouter | CLAWROUTER_API_KEY | clawrouter/anthropic/claude-sonnet-4-6 |
| Cohere | cohere | COHERE_API_KEY | cohere/command-a-plus-05-2026 |
| DeepInfra | deepinfra | DEEPINFRA_API_KEY | deepinfra/deepseek-ai/DeepSeek-V4-Flash |
| DeepSeek | deepseek | DEEPSEEK_API_KEY | deepseek/deepseek-v4-flash |
| Featherless AI | featherless | FEATHERLESS_API_KEY | featherless/Qwen/Qwen3-32B |
| GitHub Copilot | github-copilot | COPILOT_GITHUB_TOKEN / GH_TOKEN / GITHUB_TOKEN | - |
| GMI Cloud | gmi | GMI_API_KEY | gmi/google/gemini-3.1-flash-lite |
| Groq | groq | GROQ_API_KEY | groq/llama-3.3-70b-versatile |
| Hugging Face Inference | huggingface | HUGGINGFACE_HUB_TOKEN or HF_TOKEN | huggingface/deepseek-ai/DeepSeek-R1 |
| MiniMax | minimax / minimax-portal | MINIMAX_API_KEY / MINIMAX_OAUTH_TOKEN | minimax/MiniMax-M3 |
| Mistral | mistral | MISTRAL_API_KEY | mistral/mistral-large-latest |
| Moonshot | moonshot | MOONSHOT_API_KEY | moonshot/kimi-k2.6 |
| NVIDIA | nvidia | NVIDIA_API_KEY | nvidia/nvidia/nemotron-3-ultra-550b-a55b |
| NovitaAI | novita | NOVITA_API_KEY | novita/deepseek/deepseek-v3-0324 |
| Ollama Cloud | ollama-cloud | OLLAMA_API_KEY | ollama-cloud/kimi-k2.6 |
| OpenRouter | openrouter | OpenRouter OAuth or OPENROUTER_API_KEY | openrouter/auto |
| Qianfan | qianfan | QIANFAN_API_KEY | qianfan/deepseek-v3.2 |
| Tencent TokenHub | tencent-tokenhub | TOKENHUB_API_KEY | tencent-tokenhub/hy3-preview |
| Together | together | TOGETHER_API_KEY | together/meta-llama/Llama-3.3-70B-Instruct-Turbo |
| Venice | venice | VENICE_API_KEY | - |
| Vercel AI Gateway | vercel-ai-gateway | AI_GATEWAY_API_KEY | vercel-ai-gateway/anthropic/claude-opus-4.6 |
| Volcano Engine (Doubao) | volcengine / volcengine-plan | VOLCANO_ENGINE_API_KEY | volcengine-plan/ark-code-latest |
| xAI | xai | SuperGrok/X Premium OAuth or XAI_API_KEY | OAuth: xai/auto; API key: xai/grok-4.3 |
| Xiaomi | xiaomi / xiaomi-token-plan | XIAOMI_API_KEY / XIAOMI_TOKEN_PLAN_API_KEY | xiaomi/mimo-v2.5 / xiaomi-token-plan/mimo-v2.5-pro |
Quirks worth knowing
OpenRouter
Its app-attribution headers and Anthropic cache_control markers are only attached to verified openrouter.ai routes. While DeepSeek, Moonshot, and ZAI refs qualify for OpenRouter-managed prompt caching based on cache TTL, they do not get Anthropic cache markers. Because it acts as a proxy-style OpenAI-compatible path, native-OpenAI-only shaping is bypassed, covering serviceTier, Responses store, prompt-cache hints, and OpenAI reasoning-compat. Gemini-backed refs undergo only proxy-Gemini thought-signature sanitation.
Kilo Gateway
Gemini-backed refs use the same proxy-Gemini sanitation route; kilocode/kilo-auto/balanced and other refs that do not support proxy reasoning skip proxy reasoning injection.
MiniMax
During API-key onboarding, explicit M3 and M2.7 chat model definitions are written; image understanding remains on the plugin-owned MiniMax-VL-01 media provider.
NVIDIA
Model ids are namespaced under nvidia/<vendor>/<model> (for instance nvidia/nvidia/nemotron-...); pickers keep the literal <provider>/<model-id> composition intact, while the canonical key transmitted to the API stays single-prefixed.
xAI
Uses the xAI Responses path. The recommended path is SuperGrok/X Premium OAuth; fresh setup selects xai/auto, which follows xAI's authenticated default model without an OpenClaw update. Existing concrete model ids stay pinned. API keys still work via XAI_API_KEY or plugin config and keep grok-4.3 as the regional-safe setup default. Grok web_search reuses the same auth profile before API-key fallback. Older /fast and params.fastMode: true configurations still resolve through xAI's Grok 4.3 compatibility redirects, but new configurations should select a current model directly. tool_stream defaults on; disable via agents.defaults.models["xai/<model>"].params.tool_stream=false.
Providers via models.providers (custom/base URL)
Use models.providers (or models.json) to add custom providers or OpenAI/Anthropic-compatible proxies.
Many of the bundled provider plugins below already publish a default catalog. Use explicit models.providers.<id> entries only when you want to override the default base URL, headers, or model list.
Bundled and catalog-known routes take their compat capabilities from the owning provider plugin. A config compat block is for a custom provider/model or a different api/baseUrl route whose endpoint contract you have verified; see the custom-provider capability guide. Doctor removes legacy values that merely repeat the catalog and leaves divergent values visible for operator review.
Gateway model capability checks also read explicit models.providers.<id>.models[] metadata. If a custom or proxy model accepts images, set input: ["text", "image"] on that model so WebChat and node-origin attachment paths pass images as native model inputs instead of text-only media refs.
agents.defaults.models["provider/model"] controls aliases and per-model metadata for agents. It neither restricts overrides nor registers a new runtime model by itself. For custom provider models, also add models.providers.<provider>.models[] with at least the matching id; use agents.defaults.modelPolicy.allow separately when you want an override restriction.
Moonshot AI (Kimi)
Install @openclaw/moonshot-provider before onboarding. Add an explicit models.providers.moonshot entry only when you need to override the base URL or model metadata:
- Provider:
moonshot - Auth:
MOONSHOT_API_KEY - Example model:
moonshot/kimi-k3 - CLI:
openclaw onboard --auth-choice moonshot-api-keyoropenclaw onboard --auth-choice moonshot-api-key-cn
Kimi model IDs:
moonshot/kimi-k2.6moonshot/kimi-k3moonshot/kimi-k2.7-codemoonshot/kimi-k2.7-code-highspeedmoonshot/kimi-k2.5
{
agents: {
defaults: { model: { primary: "moonshot/kimi-k2.6" } },
},
models: {
mode: "merge",
providers: {
moonshot: {
baseUrl: "https://api.moonshot.ai/v1",
apiKey: "${MOONSHOT_API_KEY}",
api: "openai-completions",
models: [{ id: "kimi-k2.6", name: "Kimi K2.6" }],
},
},
},
}
See Moonshot AI (Kimi + Kimi Coding) for the full setup guide.
Kimi Coding
Kimi Coding uses Moonshot AI's Anthropic-compatible endpoint:
- Provider:
kimi - Auth:
KIMI_API_KEY - Kimi K3:
kimi/k3(up to 1M, tier-gated) orkimi/k3-256k(256K, lower quota use) - Kimi Code:
kimi/kimi-for-coding - Kimi Code HighSpeed:
kimi/kimi-for-coding-highspeed
{
env: { vars: { KIMI_API_KEY: "sk-..." } },
agents: {
defaults: { model: { primary: "kimi/kimi-for-coding" } },
},
}
Kimi K3 uses adaptive thinking. --thinking minimal|low selects low effort,
--thinking medium|high|adaptive selects high effort, and --thinking xhigh|max
selects max effort. Catalog pricing is $3/MTok input, $15/MTok output, and
$0.30/MTok cache reads. Legacy kimi/kimi-code and kimi/k2p5 remain
accepted as compatibility model ids and normalize to Kimi's stable API model
id; the previously published kimi/k3[1m] ref normalizes to kimi/k3 for
existing configs.
Volcano Engine (Doubao)
Volcano Engine (火山引擎) provides access to Doubao and other models in China.
- Provider:
volcengine(coding:volcengine-plan) - Auth:
VOLCANO_ENGINE_API_KEY - Example model:
volcengine-plan/ark-code-latest - CLI:
openclaw onboard --auth-choice volcengine-api-key
{
agents: {
defaults: { model: { primary: "volcengine-plan/ark-code-latest" } },
},
}
The onboarding flow starts you on the coding surface, yet the full volcengine/* catalog gets registered simultaneously.
When you're in onboarding or the configure model pickers, the Volcengine auth option pulls in both volcengine/* and volcengine-plan/* entries. If those models haven't been loaded yet, OpenClaw switches to the complete catalog rather than presenting a picker limited to the provider with nothing in it.
Standard models
volcengine/doubao-seed-1-8-251228(Doubao Seed 1.8)volcengine/doubao-seed-code-preview-251028volcengine/kimi-k2-5-260127(Kimi K2.5)volcengine/glm-4-7-251222(GLM 4.7)volcengine/deepseek-v3-2-251201(DeepSeek V3.2)
Coding models (volcengine-plan)
volcengine-plan/ark-code-latestvolcengine-plan/doubao-seed-code
BytePlus (International)
For users outside China, BytePlus ARK exposes the identical model set as Volcano Engine.
- Plugin:
@openclaw/byteplus-provider - Provider:
byteplus(coding:byteplus-plan) - Auth:
BYTEPLUS_API_KEY - Example model:
byteplus-plan/ark-code-latest - CLI:
openclaw onboard --auth-choice byteplus-api-key
Get the official plugin installed, then restart the Gateway:
openclaw plugins install @openclaw/byteplus-provider
openclaw gateway restart
{
agents: {
defaults: { model: { primary: "byteplus-plan/ark-code-latest" } },
},
}
Onboarding begins on the coding surface, but the general byteplus/* catalog is registered at the same time.
In onboarding/configure model pickers, the BytePlus auth choice prefers both byteplus/* and byteplus-plan/* rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.
Standard models
byteplus/seed-1-8-251228(Seed 1.8)byteplus/kimi-k2-5-260127(Kimi K2.5)byteplus/glm-4-7-251222(GLM 4.7)
Coding models (byteplus-plan)
byteplus-plan/ark-code-latestbyteplus-plan/kimi-k2.5byteplus-plan/glm-4.7
Synthetic
Synthetic offers Anthropic-compatible models through the synthetic provider:
- Provider:
synthetic - Auth:
SYNTHETIC_API_KEY - Example model:
synthetic/hf:MiniMaxAI/MiniMax-M3 - CLI:
openclaw onboard --auth-choice synthetic-api-key
{
agents: {
defaults: { model: { primary: "synthetic/hf:MiniMaxAI/MiniMax-M3" } },
},
models: {
mode: "merge",
providers: {
synthetic: {
baseUrl: "https://api.synthetic.new/anthropic",
apiKey: "${SYNTHETIC_API_KEY}",
api: "anthropic-messages",
models: [{ id: "hf:MiniMaxAI/MiniMax-M3", name: "MiniMax M3" }],
},
},
},
}
MiniMax
Because MiniMax relies on custom endpoints, you set it up via models.providers:
- MiniMax OAuth (Global):
--auth-choice minimax-global-oauth - MiniMax OAuth (CN):
--auth-choice minimax-cn-oauth - MiniMax API key (Global):
--auth-choice minimax-global-api - MiniMax API key (CN):
--auth-choice minimax-cn-api - Auth:
MINIMAX_API_KEYforminimax;MINIMAX_OAUTH_TOKENorMINIMAX_API_KEYforminimax-portal
Head to /providers/minimax for setup instructions, available models, and configuration examples.
Note
On MiniMax's Anthropic-compatible streaming path, OpenClaw turns off thinking by default for the M2.x family unless you explicitly enable it; MiniMax-M3 (and M3.x) defaults to the provider's omitted/adaptive thinking path.
/fast onrewritesMiniMax-M2.7toMiniMax-M2.7-highspeed.
Capability split owned by the plugin:
- Text and chat defaults remain on
minimax/MiniMax-M3 - Image generation relies on
minimax/image-01orminimax-portal/image-01 - Image understanding is handled by the plugin-owned
MiniMax-VL-01across both MiniMax authentication routes - Web search continues to use provider id
minimax
llama.cpp
The bundled llama-cpp plugin offers a single local text provider with two configuration options:
- Managed local server handles the installation and oversight of a verified llama-server along with local GGUF files.
- Existing llama-server links to a server you run yourself and pulls its model list from there.
For either route, install the plugin just once:
openclaw plugins install @openclaw/llama-cpp-provider
Both approaches depend on llama-cpp/<model> references. Check llama.cpp for setup,
discovery, authentication, and managed local embeddings.
LM Studio
LM Studio comes as a bundled provider plugin that taps into the native API:
- Provider:
lmstudio - Auth:
LM_API_TOKEN - Default inference base URL:
http://localhost:1234/v1
After that, assign a model (swap in one of the IDs from http://localhost:1234/api/v1/models):
{
agents: {
defaults: { model: { primary: "lmstudio/openai/gpt-oss-20b" } },
},
}
OpenClaw relies on LM Studio's built-in /api/v1/models and /api/v1/models/load for discovery and auto-load, with /v1/chat/completions as the default for inference. To let LM Studio JIT loading, TTL, and auto-evict manage the model lifecycle, enable models.providers.lmstudio.params.preload: false. Refer to /providers/lmstudio for setup and troubleshooting.
Ollama
Ollama arrives as a bundled provider plugin that uses Ollama's native API:
- Provider:
ollama - Auth: None needed (local server)
- Example model:
ollama/llama3.3 - Installation: https://ollama.com/download
# Install Ollama, then pull a model:
ollama pull llama3.3
{
agents: {
defaults: { model: { primary: "ollama/llama3.3" } },
},
}
Ollama is found locally at http://127.0.0.1:11434 once you opt in with OLLAMA_API_KEY, and the bundled provider plugin places Ollama straight into openclaw onboard and the model picker. See /providers/ollama for onboarding, cloud/local mode, and custom configuration.
vLLM
vLLM ships as a bundled provider plugin for local or self-hosted OpenAI-compatible servers:
- Provider:
vllm - Auth: Optional (depends on your server)
- Default base URL:
http://127.0.0.1:8000/v1
To enable local auto-discovery (any value is fine if your server skips auth enforcement):
export VLLM_API_KEY="vllm-local"
Then set a model (replace with one of the IDs returned by /v1/models):
{
agents: {
defaults: { model: { primary: "vllm/your-model-id" } },
},
}
See /providers/vllm for details.
SGLang
SGLang is included as a bundled provider plugin for fast self-hosted OpenAI-compatible servers:
- Provider:
sglang - Auth: Optional (depends on your server)
- Default base URL:
http://127.0.0.1:30000/v1
To opt in to local auto-discovery (any value works when your server does not enforce auth):
export SGLANG_API_KEY="sglang-local"
Then set a model (replace with one of the IDs returned by /v1/models):
{
agents: {
defaults: { model: { primary: "sglang/your-model-id" } },
},
}
See /providers/sglang for details.
Local proxies (LM Studio, vLLM, LiteLLM, etc.)
Example (OpenAI-compatible):
{
agents: {
defaults: {
model: { primary: "lmstudio/my-local-model" },
models: { "lmstudio/my-local-model": { alias: "Local" } },
},
},
models: {
providers: {
lmstudio: {
baseUrl: "http://localhost:1234/v1",
apiKey: "${LM_API_TOKEN}",
api: "openai-completions",
timeoutSeconds: 300,
models: [
{
id: "my-local-model",
name: "Local Model",
reasoning: false,
input: ["text"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
contextWindow: 200000,
maxTokens: 8192,
},
],
},
},
},
}
Default optional fields
For custom providers, reasoning, input, cost, contextWindow, and maxTokens are all optional. When they are left out, OpenClaw falls back to:
reasoning: falseinput: ["text"]cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }maxTokens: 8192
An omitted contextWindow stays unset so authored native-window metadata remains clear. If neither discovery nor per-model context metadata is present, context-budget callers use the standard 200000-token fallback.
Recommended: set explicit values that align with your proxy or model limits.
Proxy-route shaping rules
- When
api: "openai-completions"targets a non-native endpoint, meaning any non-emptybaseUrlwhose host differs fromapi.openai.com, OpenClaw enforcescompat.supportsDeveloperRole: falseto prevent provider 400 errors caused by unsupporteddeveloperroles. - Proxy-style OpenAI-compatible routes bypass native OpenAI-only request shaping entirely: no
service_tier, no Responsesstore, no Completionsstore, no prompt-cache hints, no OpenAI reasoning-compat payload shaping, and no hidden OpenClaw attribution headers. - For OpenAI-compatible Completions proxies requiring vendor-specific fields, configure
agents.defaults.models["provider/model"].params.extra_body(orextraBody) to inject extra JSON into the outgoing request body. - To control vLLM chat templates, set
agents.defaults.models["provider/model"].params.chat_template_kwargs. The bundled vLLM plugin automatically sendsenable_thinking: falseandforce_nonempty_content: trueforvllm/nemotron-3-*when the session thinking level is disabled. - For slow local models or remote LAN/tailnet hosts, set
models.providers.<id>.timeoutSeconds. This extends provider model HTTP request handling, covering connect, headers, body streaming, and the total guarded-fetch abort, without raising the overall agent runtime timeout. Ifagents.defaults.timeoutSecondsor a run-specific timeout is lower, increase that ceiling as well; provider timeouts cannot extend the entire run. - Model provider HTTP calls accept Surge, Clash, and sing-box fake-IP DNS answers in
198.18.0.0/15andfc00::/7only for the configured providerbaseUrlhostname. Custom/local provider endpoints also trust that exact configuredscheme://host:portorigin for guarded model requests, including loopback, LAN, and tailnet hosts. This is not a new config option; thebaseUrlyou configure extends the request policy only for that origin. Fake-IP hostname allowance and exact-origin trust operate independently. Other private, loopback, link-local, metadata, local-use NAT64 (64:ff9b:1::/48) destinations, and different ports still require an explicitmodels.providers.<id>.request.allowPrivateNetwork: trueopt-in. Setmodels.providers.<id>.request.allowPrivateNetwork: falseto disable the exact-origin trust. - If
baseUrlis empty or omitted, OpenClaw retains the default OpenAI behavior, which resolves toapi.openai.com. - For safety, an explicit
compat.supportsDeveloperRole: trueis still overridden on non-nativeopenai-completionsendpoints. - For
api: "anthropic-messages"on non-direct endpoints, meaning any provider other than canonicalanthropic, or a custommodels.providers.anthropic.baseUrlwhose host is not a publicapi.anthropic.comendpoint, OpenClaw suppresses implicit Anthropic beta headers such asclaude-code-20250219,interleaved-thinking-2025-05-14, and OAuth markers, so custom Anthropic-compatible proxies do not reject unsupported beta flags. Setmodels.providers.<id>.headers["anthropic-beta"]explicitly if your proxy needs specific beta features.
CLI examples
openclaw onboard --auth-choice opencode-zen
openclaw models set opencode/claude-opus-4-6
openclaw models list
See also: Configuration for full configuration examples.
Related
- Configuration reference - model config keys
- Model failover - fallback chains and retry behavior
- Models - model configuration and aliases
- Providers - per-provider setup guides