Gateway Restart Recovery: What Survives and Resumes
Learn what persists across gateway restarts or crashes, including conversations, scheduled jobs, and queued messages, and how interrupted work resumes automatically.
Read this when
- You want to know whether restarting the gateway loses in-progress agent work
- An agent run was interrupted by a restart, crash, or config reload
- You are debugging automatic session recovery after the gateway comes back up
Restarting the gateway does not wipe out agent state. Conversations, transcripts, scheduled jobs, background task records, and queued outbound messages all persist on disk, and any work interrupted mid-turn is detected and resumed automatically once the gateway is back online. Recovery is enabled by default and typically requires no manual action. However, exhausted infrastructure retries or a missing durable message-action authority claim may quarantine a single session until you inspect or replace it.
This page covers what survives a restart, how interrupted work gets detected, and what the automatic resume entails.
What survives a restart
| State | Storage | Behavior across restart |
|---|---|---|
| Conversation history | Per-agent SQLite database | Untouched; sessions continue from the stored transcript |
| Interrupted main-session turn | Per-agent SQLite session row and transcript | Automatically resumed or reconciled a few seconds after startup |
| Subagent runs | SQLite (shared state database) | Registry restored on boot; interrupted runs resumed |
| Background tasks | SQLite (shared state database) | Reconciled on boot; orphaned runs recovered or marked lost |
| Queued outbound deliveries | SQLite delivery queue | Drained after restart; undelivered replies are retried |
| Scheduled (cron) jobs | SQLite cron store | Schedules persist; the scheduler re-arms on boot |
| Restart continuation | SQLite restart sentinel | One-shot follow-up dispatched to the session that asked for the restart |
| Gateway terminal PTYs | Process memory | End with the old process; terminal sessions are not recovered |
Pending delivery rows drain or retry after restart. Failed rows discard their payload; only reusable or crash-ambiguous owners keep a minimal bounded or permanent receipt that prevents duplicate delivery.
Graceful restarts drain first
A requested restart (openclaw gateway restart, a config change that requires a restart, or a gateway update) does not terminate in-flight work immediately. The gateway halts new work intake, then waits for active agent turns and background tasks to complete, up to a drain budget (5 minutes by default). Consequently, most restarts interrupt nothing at all.
Only work that cannot finish within the drain budget (or any run interrupted by a forced restart or a crash) is aborted, and before that happens, each affected session is flagged for recovery.
Host sleep and process freezes
When a gateway host wakes from sleep, a virtual machine resumes, or the process continues after a long pause, the gateway detects the freeze within about 30 seconds. It restarts channel connections and refreshes cached health and presence so clients do not wait for stale sockets or snapshots to expire.
The macOS app and Linux companion cooperate with a local gateway by preparing a short suspension lease before the host sleeps and resuming it after wake. Remote gateways are not suspended when the app host sleeps. A deliberate suspension through gateway.suspend.* keeps recovery deferred until the controller resumes the gateway.
How interrupted work is detected
Three complementary mechanisms mark sessions whose turn did not finish:
- At turn admission: for an ordinary text turn on an existing main session, the gateway appends the user message, marks the session running, and records its recovery delivery claim in one SQLite transaction before model or
before_agent_replyhook execution. Control UI does this before returning thestartedacknowledgement; channel dispatch does it when the prepared turn adopts the agent run. Commands, attachments, per-turn overrides, pending deliveries, prior abort hints, plugin-owned sessions, and turns with execution hooks keep their specialized admission paths. If abefore_agent_replyhook is installed, admission records enough phase state to distinguish a completed silent result from an ambiguous side-effect window. Recovery dispatches an ordinary user-triggered agent turn, so the currently loadedbefore_agent_replyhooks run under their normal trigger rules. Ambiguous prior hook outcomes resume with restart-safe tools rather than replaying unrestricted side effects. - At shutdown: during the restart drain, every session with an active run is stamped with a recovery marker in the session store before the run is aborted.
- At startup: the gateway scans session stores for sessions that still claim to be running but have no live owner in the new process. This catches hard crashes and kills where no shutdown code ran. Stale transcript lock files are cleaned up at the same time.
Automatic resume
A few seconds after startup, the gateway re-dispatches each marked session with a synthetic system message telling the agent its previous turn was interrupted by a restart and to continue from the existing transcript. If a final reply had already been produced but not delivered, its text is included so the agent can deliver it instead of redoing the work.
Startup reconciliation retries transient failures up to three times with exponential backoff. Separately, each interrupted main-session cycle has a durable budget of three charged automatic dispatch attempts, retained across gateway restarts. OpenClaw charges an attempt before dispatch, refunds it when the gateway explicitly rejects the request before acceptance, and retains the charge when a post-dispatch result is uncertain to avoid replaying work. Foreground work that already owns the session keeps automatic recovery out until that work settles.
After the durable budget is exhausted, the session is tombstoned instead of looping forever. Inspect the failed session and use /new or /reset to start a replacement. openclaw doctor --fix can repair a stale aborted flag that conflicts with a tombstone, but it does not re-enable that recovery cycle.
Every retry reuses one durable dispatch identifier, so an ambiguous connection failure cannot start the same recovery twice. Completed Control UI turns also retain bounded durable idempotency tombstones, allowing a reconnecting outbox to retire them without re-executing the request.
Message-tool-only replies use a second durable correlation. Before a terminal same-conversation send reaches the channel, the gateway records an unresolved delivery intent on the exact session and source turn. A confirmed provider success resolves it to a durable delivered receipt; a confirmed failure clears it. Recovery completes a delivered receipt without rerunning tools. If a crash leaves the provider outcome unknown, recovery resumes with restart-safe tools so the model can inspect and report the ambiguity without replaying the external effect.
The delivered reply is also mirrored into the transcript with its source message ID. Terminal mirrors use a distinct receipt key, so a progress send with the same provider idempotency key cannot mask the terminal marker. Progress sends and receipts from older turns cannot complete the current turn. Only durable channel-ingress claims can restore message-action authority. A resumed run keeps the original source-delivery mode and source correlation, including requester identity and any same-channel/thread restriction, so the same receipt remains authoritative even if another restart happens during recovery. A message-tool-only turn without reconstructable channel authority is tombstoned because OpenClaw cannot safely mint message-action authority without the original channel-ingress claim. The terminal notice directs the user to start a replacement with /new or /reset.
Before resuming, the gateway classifies the transcript tail to choose the tool restriction for the continuation. An aborted turn is the interruption itself, so it resumes on a best-effort basis whatever abort detail the provider or worker recorded with it: partial streamed text stays in the transcript and the continuation picks up from the message beneath it, while a tool call left dangling is dropped from the next provider payload and restricted to restart-safe tools unless it is audited replay-safe. Provider failures, completed assistant tails, empty transcripts, and stale pending approvals also continue from the existing transcript. States with ambiguous side effects use restart-safe tools; otherwise the model decides what completed and what remains and can report any uncertainty to the user.
OpenClaw is also able to rebuild read-only Code Mode work that got interrupted. Runs marked this way are flagged as restart-safe, and any catalog or namespace tool calls that would produce side effects are refused before they run. When a restart lands on the wait control, the fresh gateway rebuilds the turn from its transcript and forces the rebuilt execution to stay restart-safe, even if the model drops or clears that flag. The host applies a filter across the entire reconstructed turn, permitting only audited read-only core tools and explicitly replay-safe plugin tools, and this holds even when Code Mode gets turned off after the restart. A Code Mode checkpoint that is non-replay-safe or unmatched still resumes for model reconciliation, but Code Mode controls are absent and the restart-safe tool restriction applies.
Subagents
Because subagent runs live in the shared SQLite state database, the subagent registry persists across process restarts. At boot the registry comes back and interrupted subagent sessions resume with their original task context. Two safety valves are in place:
- Runs interrupted more than 2 hours ago are finalized rather than resumed, so a gateway that was offline overnight does not bring stale work back to life.
- A session that keeps failing to recover is tombstoned as wedged, preventing recovery from looping indefinitely.
Background tasks
The background task registry is backed by SQLite and gets reconciled at boot and on a recurring schedule: durable outcomes recorded by completed runs are recovered, and runs whose owning process vanished are marked lost after a grace period instead of hanging forever.
Agent-requested restarts
When the agent itself initiates a restart (applying a config change, updating the gateway, or an explicit restart request), a restart sentinel is written to SQLite before the process exits. After boot the gateway reports the outcome back to the originating chat and fires a one-shot continuation turn so the agent resumes exactly where it stopped, on the same channel and thread.
For restart handling, the sentinel's typed SQLite columns are authoritative; its payload_json value is only a replay/debug shadow. Runtime reads, writes, and clears SQLite state with no file fallback. During the storage cutover, a bounded state migration runs at startup and through Doctor to preserve a validated restart-sentinel.json left by the older process after an update. The migration verifies the typed row and deletes the source file before normal restart handling proceeds.
Safety valves and observability
-
Crash-loop breaker: 3 unclean boots within 5 minutes trip a breaker that suppresses auto-start side services on the next boot, so a crashing gateway does not amplify itself. A continuously stable safe-mode gateway rechecks the breaker after the full unclean-boot window drains and then resumes deferred channel auto-start without requiring another gateway restart.
When the breaker is tripped, the control plane still starts, but channel plugins (and other auto-started side services) stay down until an operator manually overrides the suppression or the full window drains with no unclean boots. Recovery preserves channels that an operator manually stopped and any separate development-mode suppression. Gateway logs look like:
channel autostart suppressed by crash-loop breaker; refusing automatic start for <channel>… Start a channel manually with: openclaw gateway call channels.start --params '{"channel":"<id>"}'Operator recovery SOP:
- Confirm the gateway process is up (
openclaw gateway status/ LaunchAgent or systemd unit still running). A “channel disconnected” symptom often means suppressed autostart, not a dead gateway. - Inspect channel state:
openclaw channels status(add--probewhen useful). Look for stopped / not connected accounts while the gateway itself is healthy. - Fix the root cause of the unclean boots (bad config, plugin crash on start, missing secrets) before forcing channels back up.
- Manually start a channel while suppression is active:
openclaw gateway call channels.start --params '{"channel":"<id>"}' # optional: {"channel":"<id>","accountId":"<account>"}channels.startis a manual override; it does not disable the breaker for other channels.- Or leave the healthy gateway running until the full unclean-boot window
drains. The same process logs that the restart-loop breaker recovered and
starts the deferred configured channels.
If that message does not appear after the window plus one health-monitor
interval, inspect the gateway logs and run
openclaw doctorbefore restarting.
See also Gateway (safe mode paragraph) for the same control-plane vs channel-autostart split.
- Confirm the gateway process is up (
-
Main-session attempt budget: three charged automatic dispatch attempts per interrupted cycle; exhaustion tombstones that session until it is inspected and replaced.
-
Metrics: recovery activity is exported via Prometheus as
openclaw_session_recovery_totalandopenclaw_session_recovery_age_seconds. -
Logs: recovery decisions are logged under the
main-session-restart-recoveryandsubagent-interrupted-resumesubsystems. -
Reply hooks: resumed turns run currently loaded
before_agent_replyhooks under the normal user-trigger rules. Automatically delivered replies also run the normalreply_payload_sendinghook before channel delivery, with the recovered session, run, account, and conversation context.
What is not resumed
- Sessions excluded from main-session recovery because another owner already handles them: subagent sessions (subagent recovery), cron sessions (the scheduler re-runs on schedule), and ACP-managed sessions (the connected IDE or client owns the resume).
- Work that was never admitted: messages arriving during the drain window are rejected with an explicit restart error rather than silently queued into a dying process.
- Gateway terminal PTYs, including operator- and agent-owned terminals. They are process-local and end when the Gateway restarts.
- Standalone embedded turns cannot take over a main session with pending
restart recovery because they do not share the gateway's lifecycle owner.
Run the turn through the gateway or reset it there with
/newor/reset.