Model Providers: Configuration and CLI Reference

Learn how to configure LLM providers, set model refs, and manage auth via CLI. Essential for developers setting up model access and defaults.

Read this when

  • You need a provider-by-provider model setup reference
  • You want example configs or CLI onboarding commands for model providers

Reference for LLM/model providers (not chat channels like WhatsApp/Telegram). For model selection rules, see Models.

Quick rules

Model refs and CLI helpers

  • Model refs follow the provider/model format (for instance, opencode/claude-opus-4-6).
  • Aliases and model-specific settings are held in agents.defaults.models; the optional explicit override allowlist is agents.defaults.modelPolicy.allow.
  • Command-line helpers: openclaw onboard, openclaw models list, openclaw models set <provider/model>.
  • The provider-level output-token default is set by models.providers.*.maxTokens. Within each models.providers.*.models[] entry, contextWindow specifies the native window, contextTokens limits active input, and maxTokens overrides that model's output capacity.
  • For fallback rules, cooldown probes, and session-override persistence, refer to Model failover.

Adding provider auth does not change your primary model

When you add or reauth a provider, openclaw configure keeps an existing agents.defaults.model.primary intact. openclaw models auth login behaves the same way unless --set-default is supplied. Provider plugins may still suggest a default model in their auth config patch, but OpenClaw interprets that as "make this model available" when a primary model is already present, not "replace the current primary model."

To deliberately change the default model, run openclaw models set <provider/model> or openclaw models auth login --provider <id> --set-default.

OpenAI provider/runtime split

OpenAI model refs and agent runtimes are distinct concepts:

  • The canonical OpenAI provider and model are chosen via openai/<model>. A prefix alone never selects Codex.
  • When provider/model runtime policy is unset or set to auto, OpenAI may implicitly pick Codex only for an exact official HTTPS Platform Responses or ChatGPT Responses route with no authored provider request override. Valid model-scoped Fast-mode controls do not qualify as authored request params.
  • Authored Completions adapters, custom endpoints, and routes with authored request behavior remain on OpenClaw. Plaintext official HTTP endpoints are rejected.
  • Legacy Codex model refs are legacy config that doctor rewrites to openai/<model>.
  • An otherwise eligible route is explicitly kept on OpenClaw by provider/model agentRuntime.id: "openclaw". agentRuntime.id: "codex" demands Codex and fails closed when the effective route is not Codex-compatible.

Check OpenAI implicit agent runtime and Codex harness. If the provider/runtime split is unclear, start with Agent runtimes.

Plugin auto-enable follows the same boundary: an implicitly Codex-compatible effective route can enable the Codex plugin, while explicit provider/model agentRuntime.id: "codex" or legacy codex/<model> refs require it. An openai/* prefix alone does not.

Fresh OpenAI API-key and ChatGPT/Codex OAuth setup select the canonical openai/gpt-5.6-sol ref. The bare direct-API openai/gpt-5.6 alias remains supported and resolves to Sol. Existing explicit primaries, including openai/gpt-5.5, are preserved when OpenAI auth is added or refreshed. GPT-5.5 remains available through either runtime as an explicit recovery choice for accounts without GPT-5.6 access.

CLI runtimes

CLI runtimes use the same split: choose canonical model refs such as anthropic/claude-* or google/gemini-*, then set provider/model runtime policy to claude-cli or google-gemini-cli when you want a local CLI backend.

Legacy claude-cli/* and google-gemini-cli/* refs migrate back to canonical provider refs with the runtime recorded separately. Legacy codex-cli/* refs migrate to openai/* and use the Codex app-server route; OpenClaw no longer keeps a bundled Codex CLI backend.

Configure providers in the Control UI

Open Settings → Model Providers in the Control UI to add, replace, or remove provider API keys stored in models.providers.<id>.apiKey. The page identifies whether each API key comes from OpenClaw config or an environment variable without displaying the credential. Environment-provided keys remain managed by the gateway process environment.

Use Test connection to run a live provider probe and see latency or a categorized authentication, rate-limit, billing, timeout, or response error. A probe makes a real provider request and may consume a small number of tokens. OAuth and token profiles can also be logged out from the provider card.

The Default models card manages the primary model, ordered fallbacks, and utility model from the configured model catalog. Choose the models, then save them together to the existing agents.defaults.model and agents.defaults.utilityModel settings. For the utility model, Automatic leaves the setting unset and Disabled stores an empty string to turn utility routing off.

Plugin-owned provider behavior

Most provider-specific logic lives in provider plugins (registerProvider(...)) while OpenClaw keeps the generic inference loop. Plugins own onboarding, model catalogs, auth env-var mapping, transport/config normalization, tool-schema cleanup, failover classification, OAuth refresh, usage reporting, thinking/reasoning profiles, and more.

The full list of provider-SDK hooks and bundled-plugin examples lives in Provider plugins. A provider that needs a totally custom request executor is a separate, deeper extension surface.

Note

Provider-owned runner behavior lives on explicit provider hooks such as replay policy, tool-schema normalization, stream wrapping, and transport/request helpers. The legacy ProviderPlugin.capabilities static bag is compatibility-only and is no longer read by shared runner logic.

API key rotation

Key sources and priority

Configure multiple keys via:

  • OPENCLAW_LIVE_<PROVIDER>_KEY (single live override, highest priority)
  • <PROVIDER>_API_KEYS (comma or semicolon list)
  • <PROVIDER>_API_KEY (primary key)
  • <PROVIDER>_API_KEY_* (numbered list, e.g. <PROVIDER>_API_KEY_1)

For Google providers, GOOGLE_API_KEY is additionally included as a fallback. The key selection order maintains priority and removes duplicate values.

When rotation kicks in

  • Retries with the next key happen only on rate-limit responses (such as 429, rate_limit, quota, resource exhausted, Too many concurrent requests, ThrottlingException, concurrency limit reached, workers_ai ... quota limit exceeded, or recurring usage-limit notices).
  • Failures that are not rate-limit related stop immediately; no key rotation is attempted.
  • If every candidate key fails, the error from the last attempt is returned as the final result.

Official provider plugins

Official provider plugins publish their own model catalog rows. These providers do not require any models.providers model entries; enable the provider plugin, configure auth, and select a model. Use models.providers only for explicit custom providers or narrow request settings like timeouts.

OpenAI

  • Provider: openai
  • Auth: OPENAI_API_KEY
  • Optional rotation: OPENAI_API_KEYS, OPENAI_API_KEY_1, OPENAI_API_KEY_2, plus OPENCLAW_LIVE_OPENAI_KEY (single override)
  • Fresh setup default: openai/gpt-5.6-sol.
  • Example models: openai/gpt-5.6-sol, openai/gpt-5.6-terra, openai/gpt-5.6-luna, openai/gpt-5.5; the bare direct-API openai/gpt-5.6 alias remains supported.
  • Verify account/model availability with openclaw models list --provider openai if a specific install or API key behaves differently.
  • CLI: openclaw onboard --auth-choice openai-api-key
  • Direct OpenAI API-key Responses requests default to "sse".
  • Override per model via agents.defaults.models["openai/<model>"].params.transport ("sse", "websocket", "websocket-cached", or "auto"). Cached WebSockets reuse the session connection and send only new input with previous_response_id when history still matches.
  • Set an explicit OpenAI API service tier with params.serviceTier or params.service_tier; Fast mode (formerly Priority processing) uses service_tier=priority.
  • On native public OpenAI and ChatGPT/Codex Responses requests, precedence is payload/transport service_tier, then a valid explicit model param, then the fast-mode default.
  • /fast and valid params.fastMode / params.fast_mode values are shared agent-runtime controls; on direct embedded openai/* Responses requests they supply service_tier=priority only when no higher-precedence tier exists.
  • Hidden OpenClaw attribution headers (originator, version, User-Agent) apply only on native OpenAI traffic to api.openai.com, not generic OpenAI-compatible proxies
  • Native OpenAI routes also keep Responses store, prompt-cache hints, and OpenAI reasoning-compat payload shaping; proxy routes do not
  • openai/gpt-5.3-codex-spark is available only through ChatGPT/Codex OAuth; direct OpenAI API-key and Azure API-key routes reject it
{
  agents: { defaults: { model: { primary: "openai/gpt-5.6-sol" } } },
}

If the API organization does not expose GPT-5.6, set openai/gpt-5.5 explicitly. Normal onboarding and reauthentication preserve an existing explicit primary model; models auth login --set-default and models set are the intentional replacement paths.

Anthropic

  • Provider: anthropic
  • Auth: ANTHROPIC_API_KEY
  • Optional rotation: ANTHROPIC_API_KEYS, ANTHROPIC_API_KEY_1, ANTHROPIC_API_KEY_2, plus OPENCLAW_LIVE_ANTHROPIC_KEY (single override)
  • Example model: anthropic/claude-opus-5
  • CLI: openclaw onboard --auth-choice apiKey
  • Direct public Anthropic requests support the shared /fast toggle and params.fastMode, including API-key and OAuth-authenticated traffic sent to api.anthropic.com; OpenClaw maps that to Anthropic service_tier (auto vs standard_only)
  • Preferred Claude CLI config keeps the model ref canonical and selects the CLI backend separately: anthropic/claude-opus-5 with model-scoped agentRuntime.id: "claude-cli". Legacy claude-cli/claude-opus-4-7 refs still work for compatibility.

Note

Claude CLI reuse (claude -p) is a sanctioned OpenClaw integration path. Anthropic setup-token auth remains supported, but OpenClaw prefers Claude CLI reuse when available.

{
  agents: { defaults: { model: { primary: "anthropic/claude-opus-5" } } },
}

OpenAI ChatGPT/Codex OAuth

  • Provider: openai
  • Auth: OAuth (ChatGPT)
  • Fresh native Codex app-server harness ref: openai/gpt-5.6-sol
  • Native Codex app-server harness docs: Codex harness
  • Legacy model refs: codex/gpt-*, openai-codex/gpt-*
  • Plugin boundary: openai/* loads the OpenAI plugin; explicit runtime policy or the provider-owned effective route decides whether the native Codex app-server plugin is selected.
  • CLI: openclaw onboard --auth-choice openai or openclaw models auth login --provider openai
  • OpenClaw's embedded ChatGPT Responses transport defaults to auto (WebSocket-first, SSE fallback).
  • agents.defaults.models["openai/<model>"].params.transport and params.serviceTier are authored embedded-provider request settings. They keep implicit runtime selection on OpenClaw; native Codex owns its app-server transport and service tier.
  • Valid model-scoped params.fastMode / params.fast_mode values and valid cutoff keys are portable typed agent-runtime controls. They do not count as authored provider request params and do not select a runtime. Pin agentRuntime.id: "openclaw" or agentRuntime.id: "codex" when a recipe depends on one runtime.
  • Hidden OpenClaw attribution headers (originator, version, User-Agent) are only attached on native Codex traffic to chatgpt.com/backend-api, not generic OpenAI-compatible proxies
  • The shared /fast toggle, configured defaults, and valid model-scoped Fast params resolve through one runtime-control policy. See Thinking levels for precedence.
  • OpenAI API Fast mode is premium-priced and model-specific. GPT-5.6 Sol currently costs 2× Standard token pricing, and long-context multipliers stack. ChatGPT/Codex-credit Fast mode is separate: GPT-5.6 and GPT-5.5 currently consume 2.5× Standard credits, while API-key Codex runs use API token pricing. See Fast mode, API pricing, and Codex speed.
  • The native Codex catalog can expose exact openai/gpt-5.6-sol, openai/gpt-5.6-terra, and openai/gpt-5.6-luna refs according to account access. It does not apply the direct API's bare gpt-5.6 alias client-side.
  • openai/gpt-5.5 uses the Codex catalog native contextWindow = 400000 and default runtime contextTokens = 272000; override the runtime cap with models.providers.openai.models[].contextTokens
  • Sign in with openai auth and use openai/gpt-5.6-sol for a fresh subscription-backed setup. Select openai/gpt-5.5 explicitly if that Codex workspace does not expose GPT-5.6.
  • Use provider/model agentRuntime.id: "openclaw" to keep an otherwise eligible route on the built-in runtime. With runtime unset or auto, only an exact official HTTPS Responses/ChatGPT-compatible route with no authored provider request override may select Codex implicitly.
  • Legacy Codex GPT refs are legacy state, not a live provider route. Use canonical openai/* refs for new agent config, and run openclaw doctor --fix to migrate codex/* and openai-codex/* refs while preserving their native Codex semantics with model-scoped agentRuntime.id: "codex". Existing explicit canonical openai/gpt-5.5 selections are not upgraded.
{
  plugins: { entries: { codex: { enabled: true } } },
  agents: {
    defaults: {
      model: { primary: "openai/gpt-5.6-sol" },
    },
  },
}
{
  models: {
    providers: {
      openai: {
        models: [{ id: "gpt-5.5", contextTokens: 160000 }],
      },
    },
  },
}

Other subscription-style hosted options

  • MiniMax, MiniMax Coding Plan OAuth or API key access.

  • Qwen Cloud, Qwen Cloud provider surface plus Alibaba DashScope and Coding Plan endpoint mapping.

  • Z.AI (GLM), Z.AI Coding Plan or general API endpoints.

OpenCode

  • Auth: OPENCODE_API_KEY (or OPENCODE_ZEN_API_KEY)
  • Zen runtime provider: opencode
  • Go runtime provider: opencode-go
  • Example models: opencode/claude-opus-4-6, opencode-go/kimi-k2.6
  • CLI: openclaw onboard --auth-choice opencode-zen or openclaw onboard --auth-choice opencode-go
{
  agents: { defaults: { model: { primary: "opencode/claude-opus-4-6" } } },
}

Google Gemini (API key)

  • Provider: google
  • Auth: GEMINI_API_KEY
  • Optional rotation: GEMINI_API_KEYS, GEMINI_API_KEY_1, GEMINI_API_KEY_2, GOOGLE_API_KEY fallback, and OPENCLAW_LIVE_GEMINI_KEY (single override)
  • Example models: google/gemini-3.1-pro-preview, google/gemini-3.5-flash
  • Compatibility: legacy OpenClaw config using google/gemini-3.1-flash-preview is normalized to google/gemini-3-flash-preview
  • Alias: google/gemini-3.1-pro is accepted and normalized to Google's live Gemini API id, google/gemini-3.1-pro-preview
  • CLI: openclaw onboard --auth-choice gemini-api-key
  • Thinking: /think adaptive uses Google dynamic thinking. Gemini 3/3.1 omit a fixed thinkingLevel; Gemini 2.5 sends thinkingBudget: -1.
  • Direct Gemini runs also accept agents.defaults.models["google/<model>"].params.cachedContent (or legacy cached_content) to forward a provider-native cachedContents/... handle; Gemini cache hits surface as OpenClaw cacheRead

Google Vertex and Gemini CLI runtime

  • google-vertex: managed Google Cloud access through gcloud Application Default Credentials.
  • google-gemini-cli: optional local runtime for an explicitly configured canonical google/* model.

OpenClaw does not create Gemini CLI OAuth or Antigravity OAuth profiles. Connect Google through an AI Studio API key or Vertex AI. If you explicitly choose the Gemini CLI runtime, it can use the selected Google API-key profile. Existing valid Gemini CLI OAuth profiles remain runtime-compatible, but they are not a setup or recovery route.

Gemini CLI uses stream-json by default. OpenClaw reads assistant stream messages and normalizes stats.cached into cacheRead; legacy --output-format json overrides still read reply text from response.

Z.AI (GLM)

  • Provider: zai
  • Auth: ZAI_API_KEY
  • Example model: zai/glm-5.2
  • CLI: openclaw onboard --auth-choice zai-api-key
    • Model refs use the canonical zai/* provider ID.
    • zai-api-key auto-detects the matching Z.AI endpoint; zai-coding-global, zai-coding-cn, zai-global, and zai-cn force a specific surface

Vercel AI Gateway

  • Provider: vercel-ai-gateway
  • Auth: AI_GATEWAY_API_KEY
  • Example models: vercel-ai-gateway/anthropic/claude-opus-4.6, vercel-ai-gateway/moonshotai/kimi-k2.6
  • CLI: openclaw onboard --auth-choice ai-gateway-api-key

Other bundled provider plugins

ProviderIdAuth envExample model
ArceearceeARCEEAI_API_KEY or OPENROUTER_API_KEYarcee/trinity-large-thinking
BytePlusbyteplus / byteplus-planBYTEPLUS_API_KEYbyteplus-plan/ark-code-latest
CerebrascerebrasCEREBRAS_API_KEYcerebras/zai-glm-4.7
ChuteschutesCHUTES_API_KEY or CHUTES_OAUTH_TOKENchutes/zai-org/GLM-5-TEE
ClawRouterclawrouterCLAWROUTER_API_KEYclawrouter/anthropic/claude-sonnet-4-6
CoherecohereCOHERE_API_KEYcohere/command-a-plus-05-2026
DeepInfradeepinfraDEEPINFRA_API_KEYdeepinfra/deepseek-ai/DeepSeek-V4-Flash
DeepSeekdeepseekDEEPSEEK_API_KEYdeepseek/deepseek-v4-flash
Featherless AIfeatherlessFEATHERLESS_API_KEYfeatherless/Qwen/Qwen3-32B
GitHub Copilotgithub-copilotCOPILOT_GITHUB_TOKEN / GH_TOKEN / GITHUB_TOKEN-
GMI CloudgmiGMI_API_KEYgmi/google/gemini-3.1-flash-lite
GroqgroqGROQ_API_KEYgroq/llama-3.3-70b-versatile
Hugging Face InferencehuggingfaceHUGGINGFACE_HUB_TOKEN or HF_TOKENhuggingface/deepseek-ai/DeepSeek-R1
MiniMaxminimax / minimax-portalMINIMAX_API_KEY / MINIMAX_OAUTH_TOKENminimax/MiniMax-M3
MistralmistralMISTRAL_API_KEYmistral/mistral-large-latest
MoonshotmoonshotMOONSHOT_API_KEYmoonshot/kimi-k2.6
NVIDIAnvidiaNVIDIA_API_KEYnvidia/nvidia/nemotron-3-ultra-550b-a55b
NovitaAInovitaNOVITA_API_KEYnovita/deepseek/deepseek-v3-0324
Ollama Cloudollama-cloudOLLAMA_API_KEYollama-cloud/kimi-k2.6
OpenRouteropenrouterOpenRouter OAuth or OPENROUTER_API_KEYopenrouter/auto
QianfanqianfanQIANFAN_API_KEYqianfan/deepseek-v3.2
Tencent TokenHubtencent-tokenhubTOKENHUB_API_KEYtencent-tokenhub/hy3-preview
TogethertogetherTOGETHER_API_KEYtogether/meta-llama/Llama-3.3-70B-Instruct-Turbo
VeniceveniceVENICE_API_KEY-
Vercel AI Gatewayvercel-ai-gatewayAI_GATEWAY_API_KEYvercel-ai-gateway/anthropic/claude-opus-4.6
Volcano Engine (Doubao)volcengine / volcengine-planVOLCANO_ENGINE_API_KEYvolcengine-plan/ark-code-latest
xAIxaiSuperGrok/X Premium OAuth or XAI_API_KEYOAuth: xai/auto; API key: xai/grok-4.3
Xiaomixiaomi / xiaomi-token-planXIAOMI_API_KEY / XIAOMI_TOKEN_PLAN_API_KEYxiaomi/mimo-v2.5 / xiaomi-token-plan/mimo-v2.5-pro

Quirks worth knowing

OpenRouter

Its app-attribution headers and Anthropic cache_control markers are only attached to verified openrouter.ai routes. While DeepSeek, Moonshot, and ZAI refs qualify for OpenRouter-managed prompt caching based on cache TTL, they do not get Anthropic cache markers. Because it acts as a proxy-style OpenAI-compatible path, native-OpenAI-only shaping is bypassed, covering serviceTier, Responses store, prompt-cache hints, and OpenAI reasoning-compat. Gemini-backed refs undergo only proxy-Gemini thought-signature sanitation.

Kilo Gateway

Gemini-backed refs use the same proxy-Gemini sanitation route; kilocode/kilo-auto/balanced and other refs that do not support proxy reasoning skip proxy reasoning injection.

MiniMax

During API-key onboarding, explicit M3 and M2.7 chat model definitions are written; image understanding remains on the plugin-owned MiniMax-VL-01 media provider.

NVIDIA

Model ids are namespaced under nvidia/<vendor>/<model> (for instance nvidia/nvidia/nemotron-...); pickers keep the literal <provider>/<model-id> composition intact, while the canonical key transmitted to the API stays single-prefixed.

xAI

Uses the xAI Responses path. The recommended path is SuperGrok/X Premium OAuth; fresh setup selects xai/auto, which follows xAI's authenticated default model without an OpenClaw update. Existing concrete model ids stay pinned. API keys still work via XAI_API_KEY or plugin config and keep grok-4.3 as the regional-safe setup default. Grok web_search reuses the same auth profile before API-key fallback. Older /fast and params.fastMode: true configurations still resolve through xAI's Grok 4.3 compatibility redirects, but new configurations should select a current model directly. tool_stream defaults on; disable via agents.defaults.models["xai/<model>"].params.tool_stream=false.

Providers via models.providers (custom/base URL)

Use models.providers (or models.json) to add custom providers or OpenAI/Anthropic-compatible proxies.

Many of the bundled provider plugins below already publish a default catalog. Use explicit models.providers.<id> entries only when you want to override the default base URL, headers, or model list.

Bundled and catalog-known routes take their compat capabilities from the owning provider plugin. A config compat block is for a custom provider/model or a different api/baseUrl route whose endpoint contract you have verified; see the custom-provider capability guide. Doctor removes legacy values that merely repeat the catalog and leaves divergent values visible for operator review.

Gateway model capability checks also read explicit models.providers.<id>.models[] metadata. If a custom or proxy model accepts images, set input: ["text", "image"] on that model so WebChat and node-origin attachment paths pass images as native model inputs instead of text-only media refs.

agents.defaults.models["provider/model"] controls aliases and per-model metadata for agents. It neither restricts overrides nor registers a new runtime model by itself. For custom provider models, also add models.providers.<provider>.models[] with at least the matching id; use agents.defaults.modelPolicy.allow separately when you want an override restriction.

Moonshot AI (Kimi)

Install @openclaw/moonshot-provider before onboarding. Add an explicit models.providers.moonshot entry only when you need to override the base URL or model metadata:

  • Provider: moonshot
  • Auth: MOONSHOT_API_KEY
  • Example model: moonshot/kimi-k3
  • CLI: openclaw onboard --auth-choice moonshot-api-key or openclaw onboard --auth-choice moonshot-api-key-cn

Kimi model IDs:

  • moonshot/kimi-k2.6
  • moonshot/kimi-k3
  • moonshot/kimi-k2.7-code
  • moonshot/kimi-k2.7-code-highspeed
  • moonshot/kimi-k2.5
{
  agents: {
    defaults: { model: { primary: "moonshot/kimi-k2.6" } },
  },
  models: {
    mode: "merge",
    providers: {
      moonshot: {
        baseUrl: "https://api.moonshot.ai/v1",
        apiKey: "${MOONSHOT_API_KEY}",
        api: "openai-completions",
        models: [{ id: "kimi-k2.6", name: "Kimi K2.6" }],
      },
    },
  },
}

See Moonshot AI (Kimi + Kimi Coding) for the full setup guide.

Kimi Coding

Kimi Coding uses Moonshot AI's Anthropic-compatible endpoint:

  • Provider: kimi
  • Auth: KIMI_API_KEY
  • Kimi K3: kimi/k3 (up to 1M, tier-gated) or kimi/k3-256k (256K, lower quota use)
  • Kimi Code: kimi/kimi-for-coding
  • Kimi Code HighSpeed: kimi/kimi-for-coding-highspeed
{
  env: { vars: { KIMI_API_KEY: "sk-..." } },
  agents: {
    defaults: { model: { primary: "kimi/kimi-for-coding" } },
  },
}

Kimi K3 uses adaptive thinking. --thinking minimal|low selects low effort, --thinking medium|high|adaptive selects high effort, and --thinking xhigh|max selects max effort. Catalog pricing is $3/MTok input, $15/MTok output, and $0.30/MTok cache reads. Legacy kimi/kimi-code and kimi/k2p5 remain accepted as compatibility model ids and normalize to Kimi's stable API model id; the previously published kimi/k3[1m] ref normalizes to kimi/k3 for existing configs.

Volcano Engine (Doubao)

Volcano Engine (火山引擎) provides access to Doubao and other models in China.

  • Provider: volcengine (coding: volcengine-plan)
  • Auth: VOLCANO_ENGINE_API_KEY
  • Example model: volcengine-plan/ark-code-latest
  • CLI: openclaw onboard --auth-choice volcengine-api-key
{
  agents: {
    defaults: { model: { primary: "volcengine-plan/ark-code-latest" } },
  },
}

The onboarding flow starts you on the coding surface, yet the full volcengine/* catalog gets registered simultaneously.

When you're in onboarding or the configure model pickers, the Volcengine auth option pulls in both volcengine/* and volcengine-plan/* entries. If those models haven't been loaded yet, OpenClaw switches to the complete catalog rather than presenting a picker limited to the provider with nothing in it.

Standard models

  • volcengine/doubao-seed-1-8-251228 (Doubao Seed 1.8)
  • volcengine/doubao-seed-code-preview-251028
  • volcengine/kimi-k2-5-260127 (Kimi K2.5)
  • volcengine/glm-4-7-251222 (GLM 4.7)
  • volcengine/deepseek-v3-2-251201 (DeepSeek V3.2)

Coding models (volcengine-plan)

  • volcengine-plan/ark-code-latest
  • volcengine-plan/doubao-seed-code

BytePlus (International)

For users outside China, BytePlus ARK exposes the identical model set as Volcano Engine.

  • Plugin: @openclaw/byteplus-provider
  • Provider: byteplus (coding: byteplus-plan)
  • Auth: BYTEPLUS_API_KEY
  • Example model: byteplus-plan/ark-code-latest
  • CLI: openclaw onboard --auth-choice byteplus-api-key

Get the official plugin installed, then restart the Gateway:

openclaw plugins install @openclaw/byteplus-provider
openclaw gateway restart
{
  agents: {
    defaults: { model: { primary: "byteplus-plan/ark-code-latest" } },
  },
}

Onboarding begins on the coding surface, but the general byteplus/* catalog is registered at the same time.

In onboarding/configure model pickers, the BytePlus auth choice prefers both byteplus/* and byteplus-plan/* rows. If those models are not loaded yet, OpenClaw falls back to the unfiltered catalog instead of showing an empty provider-scoped picker.

Standard models

  • byteplus/seed-1-8-251228 (Seed 1.8)
  • byteplus/kimi-k2-5-260127 (Kimi K2.5)
  • byteplus/glm-4-7-251222 (GLM 4.7)

Coding models (byteplus-plan)

  • byteplus-plan/ark-code-latest
  • byteplus-plan/kimi-k2.5
  • byteplus-plan/glm-4.7

Synthetic

Synthetic offers Anthropic-compatible models through the synthetic provider:

  • Provider: synthetic
  • Auth: SYNTHETIC_API_KEY
  • Example model: synthetic/hf:MiniMaxAI/MiniMax-M3
  • CLI: openclaw onboard --auth-choice synthetic-api-key
{
  agents: {
    defaults: { model: { primary: "synthetic/hf:MiniMaxAI/MiniMax-M3" } },
  },
  models: {
    mode: "merge",
    providers: {
      synthetic: {
        baseUrl: "https://api.synthetic.new/anthropic",
        apiKey: "${SYNTHETIC_API_KEY}",
        api: "anthropic-messages",
        models: [{ id: "hf:MiniMaxAI/MiniMax-M3", name: "MiniMax M3" }],
      },
    },
  },
}

MiniMax

Because MiniMax relies on custom endpoints, you set it up via models.providers:

  • MiniMax OAuth (Global): --auth-choice minimax-global-oauth
  • MiniMax OAuth (CN): --auth-choice minimax-cn-oauth
  • MiniMax API key (Global): --auth-choice minimax-global-api
  • MiniMax API key (CN): --auth-choice minimax-cn-api
  • Auth: MINIMAX_API_KEY for minimax; MINIMAX_OAUTH_TOKEN or MINIMAX_API_KEY for minimax-portal

Head to /providers/minimax for setup instructions, available models, and configuration examples.

Note

On MiniMax's Anthropic-compatible streaming path, OpenClaw turns off thinking by default for the M2.x family unless you explicitly enable it; MiniMax-M3 (and M3.x) defaults to the provider's omitted/adaptive thinking path. /fast on rewrites MiniMax-M2.7 to MiniMax-M2.7-highspeed.

Capability split owned by the plugin:

  • Text and chat defaults remain on minimax/MiniMax-M3
  • Image generation relies on minimax/image-01 or minimax-portal/image-01
  • Image understanding is handled by the plugin-owned MiniMax-VL-01 across both MiniMax authentication routes
  • Web search continues to use provider id minimax

llama.cpp

The bundled llama-cpp plugin offers a single local text provider with two configuration options:

  • Managed local server handles the installation and oversight of a verified llama-server along with local GGUF files.
  • Existing llama-server links to a server you run yourself and pulls its model list from there.

For either route, install the plugin just once:

openclaw plugins install @openclaw/llama-cpp-provider

Both approaches depend on llama-cpp/<model> references. Check llama.cpp for setup, discovery, authentication, and managed local embeddings.

LM Studio

LM Studio comes as a bundled provider plugin that taps into the native API:

  • Provider: lmstudio
  • Auth: LM_API_TOKEN
  • Default inference base URL: http://localhost:1234/v1

After that, assign a model (swap in one of the IDs from http://localhost:1234/api/v1/models):

{
  agents: {
    defaults: { model: { primary: "lmstudio/openai/gpt-oss-20b" } },
  },
}

OpenClaw relies on LM Studio's built-in /api/v1/models and /api/v1/models/load for discovery and auto-load, with /v1/chat/completions as the default for inference. To let LM Studio JIT loading, TTL, and auto-evict manage the model lifecycle, enable models.providers.lmstudio.params.preload: false. Refer to /providers/lmstudio for setup and troubleshooting.

Ollama

Ollama arrives as a bundled provider plugin that uses Ollama's native API:

# Install Ollama, then pull a model:
ollama pull llama3.3
{
  agents: {
    defaults: { model: { primary: "ollama/llama3.3" } },
  },
}

Ollama is found locally at http://127.0.0.1:11434 once you opt in with OLLAMA_API_KEY, and the bundled provider plugin places Ollama straight into openclaw onboard and the model picker. See /providers/ollama for onboarding, cloud/local mode, and custom configuration.

vLLM

vLLM ships as a bundled provider plugin for local or self-hosted OpenAI-compatible servers:

  • Provider: vllm
  • Auth: Optional (depends on your server)
  • Default base URL: http://127.0.0.1:8000/v1

To enable local auto-discovery (any value is fine if your server skips auth enforcement):

export VLLM_API_KEY="vllm-local"

Then set a model (replace with one of the IDs returned by /v1/models):

{
  agents: {
    defaults: { model: { primary: "vllm/your-model-id" } },
  },
}

See /providers/vllm for details.

SGLang

SGLang is included as a bundled provider plugin for fast self-hosted OpenAI-compatible servers:

  • Provider: sglang
  • Auth: Optional (depends on your server)
  • Default base URL: http://127.0.0.1:30000/v1

To opt in to local auto-discovery (any value works when your server does not enforce auth):

export SGLANG_API_KEY="sglang-local"

Then set a model (replace with one of the IDs returned by /v1/models):

{
  agents: {
    defaults: { model: { primary: "sglang/your-model-id" } },
  },
}

See /providers/sglang for details.

Local proxies (LM Studio, vLLM, LiteLLM, etc.)

Example (OpenAI-compatible):

{
  agents: {
    defaults: {
      model: { primary: "lmstudio/my-local-model" },
      models: { "lmstudio/my-local-model": { alias: "Local" } },
    },
  },
  models: {
    providers: {
      lmstudio: {
        baseUrl: "http://localhost:1234/v1",
        apiKey: "${LM_API_TOKEN}",
        api: "openai-completions",
        timeoutSeconds: 300,
        models: [
          {
            id: "my-local-model",
            name: "Local Model",
            reasoning: false,
            input: ["text"],
            cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
            contextWindow: 200000,
            maxTokens: 8192,
          },
        ],
      },
    },
  },
}

Default optional fields

For custom providers, reasoning, input, cost, contextWindow, and maxTokens are all optional. When they are left out, OpenClaw falls back to:

  • reasoning: false
  • input: ["text"]
  • cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }
  • maxTokens: 8192

An omitted contextWindow stays unset so authored native-window metadata remains clear. If neither discovery nor per-model context metadata is present, context-budget callers use the standard 200000-token fallback.

Recommended: set explicit values that align with your proxy or model limits.

Proxy-route shaping rules

  • When api: "openai-completions" targets a non-native endpoint, meaning any non-empty baseUrl whose host differs from api.openai.com, OpenClaw enforces compat.supportsDeveloperRole: false to prevent provider 400 errors caused by unsupported developer roles.
  • Proxy-style OpenAI-compatible routes bypass native OpenAI-only request shaping entirely: no service_tier, no Responses store, no Completions store, no prompt-cache hints, no OpenAI reasoning-compat payload shaping, and no hidden OpenClaw attribution headers.
  • For OpenAI-compatible Completions proxies requiring vendor-specific fields, configure agents.defaults.models["provider/model"].params.extra_body (or extraBody) to inject extra JSON into the outgoing request body.
  • To control vLLM chat templates, set agents.defaults.models["provider/model"].params.chat_template_kwargs. The bundled vLLM plugin automatically sends enable_thinking: false and force_nonempty_content: true for vllm/nemotron-3-* when the session thinking level is disabled.
  • For slow local models or remote LAN/tailnet hosts, set models.providers.<id>.timeoutSeconds. This extends provider model HTTP request handling, covering connect, headers, body streaming, and the total guarded-fetch abort, without raising the overall agent runtime timeout. If agents.defaults.timeoutSeconds or a run-specific timeout is lower, increase that ceiling as well; provider timeouts cannot extend the entire run.
  • Model provider HTTP calls accept Surge, Clash, and sing-box fake-IP DNS answers in 198.18.0.0/15 and fc00::/7 only for the configured provider baseUrl hostname. Custom/local provider endpoints also trust that exact configured scheme://host:port origin for guarded model requests, including loopback, LAN, and tailnet hosts. This is not a new config option; the baseUrl you configure extends the request policy only for that origin. Fake-IP hostname allowance and exact-origin trust operate independently. Other private, loopback, link-local, metadata, local-use NAT64 (64:ff9b:1::/48) destinations, and different ports still require an explicit models.providers.<id>.request.allowPrivateNetwork: true opt-in. Set models.providers.<id>.request.allowPrivateNetwork: false to disable the exact-origin trust.
  • If baseUrl is empty or omitted, OpenClaw retains the default OpenAI behavior, which resolves to api.openai.com.
  • For safety, an explicit compat.supportsDeveloperRole: true is still overridden on non-native openai-completions endpoints.
  • For api: "anthropic-messages" on non-direct endpoints, meaning any provider other than canonical anthropic, or a custom models.providers.anthropic.baseUrl whose host is not a public api.anthropic.com endpoint, OpenClaw suppresses implicit Anthropic beta headers such as claude-code-20250219, interleaved-thinking-2025-05-14, and OAuth markers, so custom Anthropic-compatible proxies do not reject unsupported beta flags. Set models.providers.<id>.headers["anthropic-beta"] explicitly if your proxy needs specific beta features.

CLI examples

openclaw onboard --auth-choice opencode-zen
openclaw models set opencode/claude-opus-4-6
openclaw models list

See also: Configuration for full configuration examples.

4,657 words · updated Aug 22, 2026