Deep Troubleshooting Runbook for Gateway, Channels, and Automation
This runbook covers advanced troubleshooting for gateway, channels, automation, nodes, and browser issues. It includes command ladders for healthy signals, post-update recovery, and split brain installs.
Read this when
- The troubleshooting hub pointed you here for deeper diagnosis
- You need stable symptom based runbook sections with exact commands
This is the deep runbook. Start at /help/troubleshooting for the fast triage flow first.
Command ladder
Run in this order:
openclaw status
openclaw gateway status
openclaw logs --follow
openclaw doctor
openclaw channels status --probe
Healthy signals:
openclaw gateway statusshowsRuntime: running,Connectivity probe: ok, and aCapability: ...line.openclaw doctorreports no blocking config/service issues.openclaw channels status --probeshows live per-account transport status and, where supported,worksoraudit ok.
After an update
Use when an update finishes but the Gateway is down, channels are empty, or model calls fail with 401s.
openclaw status --all
openclaw update status --json
openclaw gateway status --deep
openclaw doctor --fix
openclaw gateway restart
Look for:
Update restartinopenclaw status/openclaw status --all. Pending or failed handoffs include the next command to run.plugin load failed: dependency tree corrupted; run openclaw doctor --fixunder Channels: the channel config still exists, but plugin registration failed before the channel could load.- Provider 401s after re-auth:
openclaw doctor --fixchecks for stale per-agent OAuth auth shadows and removes old copies so all agents resolve the current shared profile.
Split brain installs and newer config guard
Use when a gateway service unexpectedly stops after an update, or logs show one openclaw binary is older than the version that last wrote openclaw.json.
OpenClaw stamps config writes with meta.lastTouchedVersion. Read-only commands can inspect a config written by a newer OpenClaw, but process and service mutations refuse to run from an older binary. Blocked actions: gateway service start/stop/restart/uninstall, forced service reinstall, service-mode gateway startup, and gateway --force port cleanup.
which openclaw
openclaw --version
openclaw gateway status --deep
openclaw config get meta.lastTouchedVersion
Fix PATH
Fix PATH so openclaw resolves to the newer install, then rerun the action.
Reinstall the gateway service
Reinstall the intended gateway service from the newer install:
openclaw gateway install --force
openclaw gateway restart
Remove stale wrappers
Remove stale system package or old wrapper entries that still point at an old openclaw binary.
Warning
For intentional downgrade or emergency recovery only, set
OPENCLAW_ALLOW_OLDER_BINARY_DESTRUCTIVE_ACTIONS=1for the single command. Leave it unset for normal operation.
Protocol mismatch after rollback
Use when logs keep printing protocol mismatch after a downgrade or rollback. An older Gateway is running, but a newer local client process is still reconnecting with a protocol range the older Gateway cannot speak.
openclaw --version
which -a openclaw
openclaw gateway status --deep
openclaw doctor --deep
openclaw logs --follow
Look for:
protocol mismatch ... client=... v<version> min=<n> max=<n> expected=<n>in Gateway logs.Established clients:inopenclaw gateway status --deeporGateway clientsinopenclaw doctor --deep: active TCP clients connected to the Gateway port, with PIDs and command lines when the OS allows it.- A client process whose command line points at the newer OpenClaw install or wrapper you rolled back from.
Fix:
- Stop or restart the stale OpenClaw client process shown by
gateway status --deep. - Restart apps or wrappers that embed OpenClaw: local dashboards, editors, app-server helpers, or long-running
openclaw logs --followshells. - Re-run
openclaw gateway status --deeporopenclaw doctor --deepand confirm the stale client PID is gone.
Do not make an older Gateway accept a newer incompatible protocol. Protocol bumps protect the wire contract; rollback recovery is a process/version cleanup problem.
Skill symlink skipped as path escape
Use when logs include:
Skipping escaped skill path outside its configured root: ... reason=symlink-escape
Every skill root is a containment boundary. A symlink under ~/.agents/skills, <workspace>/.agents/skills, <workspace>/skills, or ~/.openclaw/skills is skipped when its real target resolves outside that root, unless the target is explicitly trusted.
Inspect the link:
ls -l ~/.agents/skills/<name>
realpath ~/.agents/skills/<name>
openclaw config get skills.load
If the target is intentional, configure both the direct skill root and the allowed symlink target:
{
skills: {
load: {
extraDirs: ["~/Projects/manager/skills"],
allowSymlinkTargets: ["~/Projects/manager/skills"],
},
},
}
Then start a new session or wait for the skills watcher to refresh. Restart the gateway if the running process predates the config change.
Do not use broad targets such as ~, /, or a whole synced project folder. Keep allowSymlinkTargets scoped to the real skill root that contains trusted SKILL.md directories.
If Skill Workshop apply should also write through those trusted symlinked workspace skill paths, enable skills.workshop.allowSymlinkTargetWrites. Keep it disabled for read-only shared skill roots.
Related:
Anthropic 429 extra usage required for long context
Use when logs/errors include: HTTP 429: rate_limit_error: Extra usage is required for long context requests.
openclaw logs --follow
openclaw models status
openclaw config get agents.defaults.models
Look for:
- Selected Anthropic model is a GA-capable 1M Claude 4.x model (Opus 4.6/4.7/4.8, Sonnet 4.6), or the model config still carries legacy
params.context1m: true. - Current Anthropic credential is not eligible for long-context usage.
- Requests fail only on long sessions/model runs that need the 1M context path.
Fix options:
Use a standard context window
Switch to a standard-window model, or remove legacy context1m from older
model config that is not GA-capable for 1M context.
Use an eligible credential
Use an Anthropic credential that is eligible for long-context requests, or switch to an Anthropic API key.
Configure fallback models
Configure fallback models so runs continue when Anthropic long-context requests are rejected.
Related:
Upstream 403 blocked responses
Use when an upstream LLM provider returns a generic 403 such as Your request was blocked.
Do not assume this is always an OpenClaw configuration issue. The response can come from an upstream security layer such as a CDN, WAF, bot-management rule, or reverse proxy in front of an OpenAI-compatible endpoint.
openclaw status
openclaw gateway status
openclaw logs --follow
Look for:
- Multiple models under the same provider failing the same way.
- HTML or generic security text instead of a normal provider API error.
- Provider-side security events for the same request time.
- A tiny direct
curlprobe succeeding while normal SDK-shaped requests fail.
Fix the provider-side filtering first when evidence points to a WAF/CDN block. Prefer a narrowly scoped allow or skip rule for the API path OpenClaw uses, and avoid disabling protection for the whole site.
Warning
A successful minimal
curldoes not guarantee that real SDK-style requests will pass through the same upstream security layer.
Related:
Local OpenAI-compatible backend passes direct probes but agent runs fail
Use when:
curl ... /v1/modelsworks.- Tiny direct
/v1/chat/completionscalls work. - OpenClaw model runs fail only on normal agent turns.
curl http://127.0.0.1:1234/v1/models
curl http://127.0.0.1:1234/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"<id>","messages":[{"role":"user","content":"hi"}],"stream":false}'
openclaw infer model run --model <provider/model> --prompt "hi" --json
openclaw logs --follow
Look for:
- Direct tiny calls succeed, but OpenClaw runs fail only on larger prompts.
model_not_foundor 404 errors even though direct/v1/chat/completionsworks with the same bare model id.- Backend errors about
messages[].contentexpecting a string. - Intermittent
incomplete turn detected ... stopReason=stop payloads=0warnings with an OpenAI-compatible local backend. - Backend crashes that appear only with larger prompt-token counts or full agent runtime prompts.
Common signatures
model_not_foundwith a local MLX/vLLM-style server: verifybaseUrlincludes/v1,apiis"openai-completions"for/v1/chat/completionsbackends, andmodels.providers.<provider>.models[].idis the bare provider-local id. Select it with the provider prefix once, for examplemlx/mlx-community/Qwen3-30B-A3B-6bit; keep the catalog entry asmlx-community/Qwen3-30B-A3B-6bit.messages[...].content: invalid type: sequence, expected a string: backend rejects structured Chat Completions content parts. Fix: setmodels.providers.<provider>.models[].compat.requiresStringContent: true.validation.keysor allowed message keys like["role","content"]: backend rejects OpenAI-style replay metadata on Chat Completions messages. Fix: setmodels.providers.<provider>.models[].compat.strictMessageKeys: true.incomplete turn detected ... stopReason=stop payloads=0: the backend completed the Chat Completions request but returned no user-visible assistant text for that turn. OpenClaw retries replay-safe empty OpenAI-compatible turns once; persistent failures usually mean the backend is emitting empty/non-text content or suppressing final-answer text.- Direct tiny requests succeed, but OpenClaw agent runs fail with backend/model crashes (for example Gemma on some
inferrsbuilds): OpenClaw transport is likely already correct; the backend is failing on the larger agent-runtime prompt shape. - Failures shrink after disabling tools but do not disappear: tool schemas were part of the pressure, but the remaining issue is still upstream model/server capacity or a backend bug.
Fix options
- Set
compat.requiresStringContent: truefor string-only Chat Completions backends. - Set
compat.strictMessageKeys: truefor strict Chat Completions backends that only acceptroleandcontenton each message. - Set
compat.supportsTools: falsefor models/backends that cannot handle OpenClaw's tool schema surface reliably. - Lower prompt pressure where possible: smaller workspace bootstrap, shorter session history, lighter local model, or a backend with stronger long-context support.
- If tiny direct requests keep passing while OpenClaw agent turns still crash inside the backend, treat it as an upstream server/model limitation and file a repro there with the accepted payload shape.
Related:
No replies
If channels are up but nothing responds, inspect routing and policy before touching any connections.
openclaw status
openclaw channels status --probe
openclaw pairing list --channel <channel> [--account <id>]
openclaw config get channels
openclaw logs --follow
Check for:
- Pairing pending for DM senders.
- Group mention gating (
requireMention,mentionPatterns). - Channel or group allowlist mismatches.
Common indicators:
drop guild message (mention requiredmeans a group message was ignored until mentioned.pairing requestmeans the sender needs approval.blocked/allowlistmeans the sender or channel was blocked by policy.
Related:
Dashboard control UI connectivity
When the dashboard or control UI will not connect, verify the URL, authentication mode, and secure context assumptions.
openclaw gateway status
openclaw status
openclaw logs --follow
openclaw doctor
openclaw gateway status --json
Check for:
- Correct probe URL and dashboard URL.
- Auth mode or token mismatch between client and gateway.
- HTTP usage where device identity is required.
If a local browser cannot reach 127.0.0.1:18789 after an update, first recover the local Gateway service and confirm it is serving the dashboard:
openclaw gateway restart
lsof -i :18789
curl http://127.0.0.1:18789
If curl returns OpenClaw HTML, the Gateway is running and the remaining problem is likely browser cache, an old deep link, or stale tab state. Open http://127.0.0.1:18789 directly and navigate from the dashboard. If restart does not keep the service running, run openclaw gateway start and recheck openclaw gateway status.
Connect / auth signatures
device identity requiredindicates a non-secure context or missing device authentication.origin not allowedmeans the browserOriginis not ingateway.controlUi.allowedOrigins(or you are connecting from a non-loopback browser origin without an explicit allowlist).device nonce required/device nonce mismatchmeans the client is not completing the challenge-based device auth flow (connect.challenge+device.nonce).device signature invalid/device signature expiredmeans the client signed the wrong payload (or a stale timestamp) for the current handshake.AUTH_TOKEN_MISMATCHwithcanRetryWithDeviceToken=truemeans the client can perform one trusted retry with a cached device token.- That cached-token retry reuses the cached scope set stored with the paired device token. Explicit
deviceToken/ explicitscopescallers keep their requested scope set instead. AUTH_SCOPE_MISMATCHmeans the device token was recognized, but its approved scopes do not cover this connect request. Re-pair or approve the requested scope contract instead of rotating a shared gateway token.- Outside that retry path, connect auth precedence is: explicit shared token or password first, then explicit
deviceToken, then stored device token, then bootstrap token. - On the async Tailscale Serve Control UI path, failed attempts for the same
{scope, ip}are serialized before the limiter records the failure. Two bad concurrent retries from the same client can therefore surfaceretry lateron the second attempt instead of two plain mismatches. too many failed authentication attempts (retry later)from a browser-origin loopback client means repeated failures from that same normalizedOriginare locked out temporarily. Another localhost origin uses a separate bucket.- Repeated
unauthorizedafter that retry indicates shared token or device token drift. Refresh token config and re-approve or rotate the device token if needed. gateway connect failed:means the wrong host, port, or URL target was used.
Auth detail codes quick map
Use error.details.code from the failed connect response to decide the next action:
| Detail code | Meaning | Recommended action |
|---|---|---|
AUTH_TOKEN_MISSING | Client did not send a required shared token. | Paste or set the token in the client and retry. For dashboard paths: openclaw config get gateway.auth.token then paste into Control UI settings. |
AUTH_TOKEN_MISMATCH | Shared token did not match the gateway auth token. | If canRetryWithDeviceToken=true, allow one trusted retry. Cached-token retries reuse stored approved scopes. Explicit deviceToken / scopes callers keep requested scopes. If still failing, run the token drift recovery checklist. |
AUTH_DEVICE_TOKEN_MISMATCH | Cached per-device token is stale or revoked. | Rotate or re-approve the device token using the devices CLI, then reconnect. |
AUTH_SCOPE_MISMATCH | Device token is valid, but its approved role or scopes do not cover this connect request. | Re-pair the device or approve the requested scope contract. Do not treat this as shared-token drift. |
PAIRING_REQUIRED | Device identity needs approval. Check error.details.reason for not-paired, scope-upgrade, role-upgrade, or metadata-upgrade, and use requestId / remediationHint when present. | Approve the pending request: openclaw devices list then openclaw devices approve <requestId>. Scope or role upgrades use the same flow after you review the requested access. |
Note
Direct loopback backend RPCs authenticated with the shared gateway token or password should not depend on the CLI's paired-device scope baseline. If subagents or other internal calls still fail with
scope-upgrade, verify the caller is usingclient.id: "gateway-client"andclient.mode: "backend"and is not forcing an explicitdeviceIdentityor device token.
Device auth v2 migration check:
openclaw --version
openclaw doctor
openclaw gateway status
If logs show nonce or signature errors, update the connecting client and verify it:
Wait for connect.challenge
Client waits for the gateway-issued connect.challenge.
Sign the payload
Client signs the challenge-bound payload.
Send the device nonce
Client sends connect.params.device.nonce with the same challenge nonce.
If openclaw devices rotate / revoke / remove is denied unexpectedly:
- Paired-device token sessions can manage only their own device unless the caller also has
operator.admin. openclaw devices rotate --scope ...can only request operator scopes that the caller session already holds.
Related:
- Configuration (gateway auth modes)
- Control UI
- Devices
- Remote access
- Trusted proxy auth
Gateway service not running
Use when the service is installed but the process does not stay running.
openclaw gateway status
openclaw status
openclaw logs --follow
openclaw doctor
openclaw gateway status --deep # also scan system-level services
Check for:
Runtime: stoppedwith exit hints.- Service config mismatch (
Config (cli)vsConfig (service)). - Port or listener conflicts.
- Extra launchd, systemd, or schtasks installs when
--deepis used. Other gateway-like services detected (best effort)cleanup hints.
Common signatures
Gateway start blocked: set gateway.mode=localorexisting config is missing gateway.modemeans local gateway mode is not enabled, or the config file was clobbered and lostgateway.mode. Fix: setgateway.mode="local"in your config, or re-runopenclaw onboard --mode local/openclaw setupto restamp the expected local-mode config. If you are running OpenClaw via Podman, the default config path is~/.openclaw/openclaw.json.refusing to bind gateway ... without authmeans a non-loopback bind without a valid gateway auth path (token or password, or trusted-proxy where configured).another gateway instance is already listening/EADDRINUSEmeans a port conflict.Other gateway-like services detected (best effort)means stale or parallel launchd, systemd, or schtasks units exist. Most setups should keep one gateway per machine. If you do need more than one, isolate ports plus config, state, and workspace. See /gateway#multiple-gateways-same-host.System-level OpenClaw gateway service detectedfrom doctor means a systemd system unit exists while the user-level service is missing. Remove or disable the duplicate before allowing doctor to install a user service, or setOPENCLAW_SERVICE_REPAIR_POLICY=externalif the system unit is the intended supervisor.Gateway service port does not match current gateway configmeans the installed supervisor still pins the old--port. Runopenclaw doctor --fixoropenclaw gateway install --force, then restart the gateway service.
Related:
macOS gateway silently stops responding, then resumes when you touch the dashboard
Use when channels (Telegram, WhatsApp, etc.) on a macOS host go quiet for minutes to hours at a time, and the gateway appears to come back the moment you open the Control UI, SSH in, or otherwise interact with the host. There is usually no obvious symptom in openclaw status because by the time you look the gateway is alive again.
ls ~/.openclaw/logs/stability/ | tail -5
openclaw gateway stability --bundle latest
pmset -g log | grep -iE "sleep|wake|maintenance" | tail -50
launchctl print gui/$UID/ai.openclaw.gateway | grep -E "state|last exit|runs"
Check for:
- One or more
*-uncaught_exception.jsonbundles in~/.openclaw/logs/stability/witherror.codeset to a transient network code such asENETDOWN,ENETUNREACH,EHOSTUNREACH, orECONNREFUSED. pmset -g loglines likeEntering Sleep state due to 'Maintenance Sleep'oren0 driver is slow (msg: WillChangeState to 0)aligned with the crash timestamps. Power Nap or Maintenance Sleep briefly puts the Wi-Fi driver into state 0. Any outboundconnect()that lands in that window can fail withENETDOWNeven on a host that otherwise has full network connectivity.launchctl printoutput showingstate = not runningwith multiple recentrunsand an exit code, especially when the gap between crash and the next launch is on the order of an hour rather than seconds. macOS launchd applies an undocumented respawn-protection gate after a crash burst that can stop honoringKeepAlive=trueuntil an external trigger such as interactive login, dashboard connection, orlaunchctl kickstartre-arms it.
Common signatures:
- A stability bundle where
error.codeisENETDOWNor a sibling code, with the call stack pointing into NodenetlookupAndConnect/Socket.connect. OpenClaw2026.5.26and later treat these as benign transient network errors, so they no longer reach the top-level uncaught handler. If you are on an older release, upgrade first. - Long quiet periods that stop as soon as you connect to the Control UI or SSH into the host. The user-visible activity re-arms launchd's respawn gate, not anything the dashboard does to the gateway.
runscount rising throughout the day with no matchingreceived SIG*; shutting downline in~/Library/Logs/openclaw/gateway.log. Clean shutdowns log a signal; transient crashes do not.
What to do:
-
Upgrade the gateway if you are running a release before
2026.5.26. After upgrading, futureENETDOWNerrors are logged as warnings instead of terminating the process. -
Reduce maintenance sleep activity on Mac mini or desktop hosts meant to run as always-on servers:
sudo pmset -a sleep 0 disksleep 0 standby 0 powernap 0This cuts down on, but does not fully eliminate, the underlying driver flap. The system can still perform some maintenance sleeps for TCP keepalive and mDNS upkeep regardless of these flags.
-
Add a liveness watchdog so a future crash burst that gets parked by launchd is caught quickly:
# Example launchd-aware liveness check, suitable for a 5-minute cron or LaunchAgent state=$(launchctl print gui/$UID/ai.openclaw.gateway 2>/dev/null | awk -F'= ' '/state =/ {print $2; exit}') if [ "$state" != "running" ]; then launchctl kickstart -k gui/$UID/ai.openclaw.gateway fiThe goal is to externally re-arm the respawn gate.
KeepAlive=truealone is not enough on macOS after a crash burst.
Related:
macOS launchd supervisor loop with duplicate gateway/node LaunchAgents
Use this when a macOS install keeps restarting every few seconds, openclaw
health checks flap between healthy and unavailable, and channel dispatch stalls
even though the service appears to be running.
This was seen on older installs where both ai.openclaw.gateway and
ai.openclaw.node LaunchAgents were active and each injected
OPENCLAW_LAUNCHD_LABEL. In that state OpenClaw can detect launchd
supervision, try to hand restart back to launchd, and fall into a fast
EADDRINUSE/respawn loop instead of one stable gateway process.
for i in 1 2 3 4; do
ps aux | grep 'openclaw.*index.js' | grep -v grep | awk '{print $2}'
sleep 10
done
openclaw gateway status --deep
openclaw node status
launchctl print gui/$UID/ai.openclaw.gateway | grep -E 'state|last exit|runs'
tail -n 80 ~/Library/Logs/openclaw/gateway.log
Look for:
- More than one gateway PID across the 30-second sample instead of one stable process.
EADDRINUSE,another gateway instance is already listening, or repeated restart/handoff lines ingateway.log.- Both
~/Library/LaunchAgents/ai.openclaw.gateway.plistand~/Library/LaunchAgents/ai.openclaw.node.plistloaded at the same time on a host that should only run one managed gateway service.
What to do:
-
If this host should only run the Gateway service, remove the managed node service through OpenClaw. Skip this step if you actively rely on the node service for remote node features. Uninstalling it stops those features on this host:
openclaw node uninstall -
Install a persistent Gateway wrapper that clears the inherited launchd markers before starting OpenClaw. Use the supported
--wrapperoption. Do not edit the generated file under~/.openclaw/service-env/, because service reinstall, update, and doctor repair regenerate that file:mkdir -p ~/.local/bin cat >~/.local/bin/openclaw-launchd-workaround <<'EOF' #!/bin/sh set -eu unset OPENCLAW_LAUNCHD_LABEL LAUNCH_JOB_LABEL LAUNCH_JOB_NAME XPC_SERVICE_NAME || true exec openclaw "$@" EOF chmod 700 ~/.local/bin/openclaw-launchd-workaround openclaw gateway install \ --wrapper ~/.local/bin/openclaw-launchd-workaround \ --forcegateway installpersists the wrapper path across forced reinstalls, updates, and doctor repairs. -
Verify that the Gateway is stable and serving RPC, not merely listening:
openclaw gateway status --deep --require-rpc for i in 1 2 3 4; do ps aux | grep 'openclaw.*index.js' | grep -v grep | awk '{print $2}' sleep 10 doneThe PID sample should show one stable process instead of a rotating set of PIDs, and inbound channel dispatch should resume.
-
After upgrading to a release where the underlying dual-LaunchAgent loop is fixed, remove the workaround and reinstall the normal managed service:
OPENCLAW_WRAPPER= openclaw gateway install --force rm ~/.local/bin/openclaw-launchd-workaround
Related:
Gateway exits during high memory use
Use when the Gateway disappears under load, the supervisor reports an OOM-style restart, or logs mention critical memory pressure bundle written.
openclaw gateway status --deep
openclaw logs --follow
openclaw gateway stability --bundle latest
openclaw gateway diagnostics export
Look for:
Reason: diagnostic.memory.pressure.criticalin the latest stability bundle.Memory pressure:withcritical/rss_threshold,critical/heap_threshold, orcritical/rss_growth.V8 heap:values near the heap limit.Largest session files:entries such asagents/<agent>/sessions/<session>.jsonlorsessions/<session>.jsonl.- Linux cgroup memory counters when the gateway runs inside a container or memory-limited service.
Common signatures:
critical memory pressure bundle writtenappears shortly before restart. OpenClaw captured a pre-OOM stability bundle. Inspect it withopenclaw gateway stability --bundle latest.memory pressure: level=criticalappears in gateway logs. OpenClaw detected critical memory pressure and recorded the available in-process memory facts.Largest session files:points at a very large redacted transcript path. Reduce retained session history, inspect session growth, or move old transcripts out of the active store before restarting.V8 heap:used bytes are close to the heap limit. Lower prompt/session pressure or reduce concurrent work first. For a managed service, inspectGateway heap:inopenclaw gateway status. If it saysnot set, regenerate old service metadata withopenclaw gateway install --force. Ambient shellNODE_OPTIONSis intentionally ignored. Use an explicit supervisor-level heap override only after confirming the sustained workload and leaving enough native-memory headroom.Memory pressure: critical/rss_growth. Memory grew quickly inside one sampling window. Check the latest logs for a large import, runaway tool output, repeated retries, or a batch of queued agent work.- Critical memory pressure appears in logs but no bundle exists. Capture
openclaw gateway diagnostics exportafter the event for the available operational evidence.
The stability bundle is payload-free. It includes operational memory evidence and redacted relative file paths, not message text, webhook bodies, credentials, tokens, cookies, or raw session ids. Attach the diagnostics export to bug reports instead of copying raw logs.
Related:
Gateway rejected invalid config
Use when Gateway startup fails with Invalid config or hot reload logs say it skipped an invalid edit.
openclaw logs --follow
openclaw config file
openclaw config validate
openclaw doctor
Look for:
Invalid config at ...config reload skipped (invalid config): ...Config write rejected: ...- A timestamped
openclaw.json.rejected.*file beside the active config. - A timestamped
openclaw.json.clobbered.*file ifdoctor --fixrepaired a broken direct edit. - OpenClaw keeps the latest 32
.clobbered.*files for each config path and rotates older ones.
What happened
- The config did not validate during startup, hot reload, or an OpenClaw-owned write.
- Gateway startup fails closed instead of rewriting
openclaw.json. - Hot reload skips invalid external edits and keeps the current runtime config active.
- OpenClaw-owned writes reject invalid/destructive payloads before commit and save
.rejected.*. openclaw doctor --fixowns repair. It can remove non-JSON prefixes or restore the last-known-good copy while preserving the rejected payload as.clobbered.*.- When many repairs happen for one config path, OpenClaw rotates older
.clobbered.*files so the newest repaired payload is still available.
Inspect and repair
CONFIG="$(openclaw config file)"
ls -lt "$CONFIG".clobbered.* "$CONFIG".rejected.* 2>/dev/null | head
diff -u "$CONFIG" "$(ls -t "$CONFIG".clobbered.* 2>/dev/null | head -n 1)"
openclaw config validate
openclaw doctor
Common signatures
.clobbered.*exists. Doctor preserved a broken external edit while repairing the active config..rejected.*exists. An OpenClaw-owned config write failed schema or clobber checks before commit.Config write rejected:. The write tried to drop required shape, shrink the file sharply, or persist invalid config.config reload skipped (invalid config):. A direct edit failed validation and was ignored by the running Gateway.Invalid config at .... Startup failed before Gateway services booted.missing-meta-vs-last-good,gateway-mode-missing-vs-last-good, orsize-drop-vs-last-good:*. An OpenClaw-owned write was rejected because it lost fields or size compared with the last-known-good backup.Config last-known-good promotion skipped. The candidate contained redacted secret placeholders such as***.
Fix options
- Run
openclaw doctor --fixto let doctor repair prefixed/clobbered config or restore last-known-good. - Copy only the intended keys from
.clobbered.*or.rejected.*, then apply them withopenclaw config setorconfig.patch. - Run
openclaw config validatebefore restarting. - If you edit by hand, keep the full JSON5 config, not just the partial object you wanted to change.
Related:
Gateway probe warnings
Use when openclaw gateway probe reaches something, but still prints a warning block.
openclaw gateway probe
openclaw gateway probe --json
openclaw gateway probe --ssh user@gateway-host
Look for:
warnings[].codeandprimaryTargetIdin JSON output.- Whether the warning is about SSH fallback, multiple gateways, missing scopes, or unresolved auth refs.
Common signatures:
SSH tunnel failed to start; falling back to direct probes.. SSH setup failed, but the command still tried direct configured/loopback targets.multiple reachable gateway identities detected. Distinct gateways answered, or OpenClaw could not prove reachable targets are the same gateway. An SSH tunnel, proxy URL, or configured remote URL to the same gateway is treated as one gateway with multiple transports, even when transport ports differ.Read-probe diagnostics are limited by gateway scopes (missing operator.read). Connect worked, but detail RPC is scope-limited. Pair device identity or use credentials withoperator.read.Gateway accepted the WebSocket connection, but follow-up read diagnostics failed. Connect worked, but the full diagnostic RPC set timed out or failed. Treat this as a reachable Gateway with degraded diagnostics. Compareconnect.okandconnect.rpcOkin--jsonoutput.Capability: pairing-pendingorgateway closed (1008): pairing required. The gateway answered, but this client still needs pairing/approval before normal operator access.- Unresolved
gateway.auth.*/gateway.remote.*SecretRef warning text. Auth material was unavailable in this command path for the failed target.
Related:
Channel connected, messages not flowing
When the channel shows as connected but no messages are flowing, check policy, permissions, and channel specific delivery rules.
openclaw channels status --probe
openclaw pairing list --channel <channel> [--account <id>]
openclaw status --deep
openclaw logs --follow
openclaw config get channels
Look for:
- DM policy (
pairing,allowlist,open,disabled). - Group allowlist and mention requirements.
- Missing channel API permissions or scopes.
Common signatures:
mention requiredmeans the message was ignored by group mention policy.pairingor pending approval traces indicate the sender is not approved.missing_scope,not_in_channel,Forbidden,401/403point to a channel auth or permissions issue.
Related:
Cron and heartbeat delivery
If a cron job or heartbeat did not run or failed to deliver, start by checking the scheduler state, then the delivery target.
openclaw cron status
openclaw cron list
openclaw cron runs --id <jobId> --limit 20
openclaw system heartbeat last
openclaw logs --follow
Look for:
- Cron enabled and next wake time present.
- Job run history status (
ok,skipped,error). - Heartbeat skip reasons (
quiet-hours,requests-in-flight,cron-in-progress,lanes-busy,alerts-disabled,empty-heartbeat-file).
Common signatures
cron: scheduler disabled; jobs will not run automaticallymeans cron is disabled.cron: timer tick failedmeans the scheduler tick failed; check file, log, or runtime errors.heartbeat skippedwithreason=quiet-hoursmeans the job is outside active hours.heartbeat skippedwithreason=empty-heartbeat-filemeans the heartbeat monitor scratch contains only blank, comment, header, fence, or empty checklist scaffolding, so OpenClaw skips the model call.heartbeat: unknown accountIdmeans an invalid account id for the heartbeat delivery target.heartbeat skippedwithreason=dm-blockedmeans the heartbeat target resolved to a DM style destination whileagents.defaults.heartbeat.directPolicy(or a per-agent override) is set toblock.
Related:
Node paired, tool fails
When a node is paired but tools fail, isolate the foreground, permission, and approval state.
openclaw nodes status
openclaw nodes describe --node <idOrNameOrIp>
openclaw approvals get --node <idOrNameOrIp>
openclaw logs --follow
openclaw status
Look for:
- Node online with expected capabilities.
- OS permission grants for camera, mic, location, or screen.
- Exec approvals and allowlist state.
Common signatures:
NODE_BACKGROUND_UNAVAILABLEmeans the node app must be in the foreground.*_PERMISSION_REQUIREDorLOCATION_PERMISSION_REQUIREDmeans a missing OS permission.SYSTEM_RUN_DENIED: approval requiredmeans an exec approval is pending.SYSTEM_RUN_DENIED: allowlist missmeans the command is blocked by the allowlist.
Related:
Browser tool fails
Use this when browser tool actions fail even though the gateway itself is healthy.
openclaw browser status
openclaw browser start --browser-profile openclaw
openclaw browser profiles
openclaw logs --follow
openclaw doctor
Look for:
- Whether
plugins.allowis set and includesbrowser. - A valid browser executable path.
- CDP profile reachability.
- Local Chrome availability for
existing-sessionoruserprofiles.
Plugin / executable signatures
unknown command "browser"orunknown command 'browser'means the bundled browser plugin is excluded byplugins.allow.- Browser tool missing or unavailable while
browser.enabled=truemeansplugins.allowexcludesbrowser, so the plugin never loaded. Failed to start Chrome CDP on portmeans the browser process failed to launch.browser.executablePath not foundmeans the configured path is invalid.browser.cdpUrl must be http(s) or ws(s)means the configured CDP URL uses an unsupported scheme such asfile:orftp:.browser.cdpUrl has invalid portmeans the configured CDP URL has a bad or out-of-range port.Playwright is not available in this gateway build; '<feature>' is unsupported.means the current gateway install lacks the core browser runtime dependency; reinstall or update OpenClaw, then restart the gateway. ARIA snapshots and basic page screenshots can still work, but navigation, AI snapshots, CSS selector element screenshots, and PDF export remain unavailable.
Chrome MCP / existing-session signatures
Could not find DevToolsActivePort for chromemeans the Chrome MCP existing session could not attach to the selected browser data dir yet. Open the browser inspect page, enable remote debugging, keep the browser open, approve the first attach prompt, then retry. If signed in state is not required, prefer the managedopenclawprofile.No browser tabs found for profile="user"means the Chrome MCP attach profile has no open local Chrome tabs.Remote CDP for profile "<name>" is not reachablemeans the configured remote CDP endpoint is not reachable from the gateway host.Browser attachOnly is enabled ... not reachableorBrowser attachOnly is enabled and CDP websocket ... is not reachablemeans the attach only profile has no reachable target, or the HTTP endpoint answered but the CDP WebSocket still could not be opened.
Element / screenshot / upload signatures
fullPage is not supported for element screenshotsmeans the screenshot request mixed--full-pagewith--refor--element.element screenshots are not supported for existing-session profiles; use ref from snapshot.means Chrome MCP orexisting-sessionscreenshot calls must use page capture or a snapshot--ref, not CSS--element.existing-session file uploads do not support element selectors; use ref/inputRef.means Chrome MCP upload hooks need snapshot refs, not CSS selectors.existing-session file uploads currently support one file at a time.means send one upload per call on Chrome MCP profiles.existing-session dialog handling does not support timeoutMs.means dialog hooks on Chrome MCP profiles do not support timeout overrides.existing-session type does not support timeoutMs overrides.means omittimeoutMsforact:typeonprofile="user"or Chrome MCP existing session profiles, or use a managed or CDP browser profile when a custom timeout is required.response body is not supported for existing-session profiles yet.meansresponsebodystill requires a managed browser or raw CDP profile.- Stale viewport, dark mode, locale, or offline overrides on attach only or remote CDP profiles means run
openclaw browser stop --browser-profile <name>to close the active control session and release Playwright or CDP emulation state without restarting the whole gateway.
Related:
If you upgraded and something suddenly broke
Most post upgrade breakage comes from config drift or stricter defaults now being enforced.
1. Auth and URL override behavior changed
openclaw gateway status
openclaw config get gateway.mode
openclaw config get gateway.remote.url
openclaw config get gateway.auth.mode
What to check:
- If
gateway.mode=remote, CLI calls may be targeting remote while your local service is fine. - Explicit
--urlcalls do not fall back to stored credentials.
Common signatures:
gateway connect failed:means the wrong URL target.unauthorizedmeans the endpoint is reachable but auth is wrong.
2. Bind and auth guardrails are stricter
openclaw config get gateway.bind
openclaw config get gateway.auth.mode
openclaw config get gateway.auth.token
openclaw gateway status
openclaw logs --follow
What to check:
- Non-loopback binds (
lan,tailnet,custom) need a valid gateway auth path: shared token or password auth, or a correctly configured non-loopbacktrusted-proxydeployment. - Old keys like
gateway.tokendo not replacegateway.auth.token.
Common signatures:
refusing to bind gateway ... without authmeans a non-loopback bind without a valid gateway auth path.Connectivity probe: failedwhile runtime is running means the gateway is alive but inaccessible with the current auth or URL.
3. Pairing and device identity state changed
openclaw devices list
openclaw pairing list --channel <channel> [--account <id>]
openclaw logs --follow
openclaw doctor
What to check:
- Pending device approvals for dashboard or nodes.
- Pending DM pairing approvals after policy or identity changes.
Common signatures:
device identity requiredmeans device auth is not satisfied.pairing requiredmeans the sender or device must be approved.
If the service config and runtime still disagree after checks, reinstall service metadata from the same profile or state directory:
openclaw gateway install --force
openclaw gateway restart
Related: