Transcript Hygiene: Provider-Specific Sanitization and Repair Rules

This reference explains how OpenClaw applies provider-specific fixes to transcripts before each run. It covers runtime prompt context, tool call sanitization, turn validation, and cleanup rules for developers managing session integrity.

Read this when

  • You are debugging provider request rejections tied to transcript shape
  • You are changing transcript sanitization or tool-call repair logic
  • You are investigating tool-call id mismatches across providers

Before each run (when building model context), OpenClaw applies provider-specific fixes to transcripts. Most of these adjustments happen in memory and exist solely to meet strict provider requirements. A separate session file repair pass may also rewrite stored JSONL before the session loads, but only for malformed lines or persisted turns that are not valid durable records. Delivered assistant replies are saved to disk; provider-specific assistant prefill stripping occurs only while building outbound payloads.

When a repair takes place, the original file is written to a temporary *.bak-<pid>-<ts> sibling before the atomic replacement, then deleted once the replacement succeeds. The backup is kept only if cleanup itself fails, in which case the path is reported back.

Scope covers:

  • Runtime only prompt context staying out of user visible transcript turns
  • Tool call id sanitization
  • Tool call input validation
  • Tool result pairing repair
  • Turn validation and ordering
  • Thought signature cleanup
  • Thinking signature cleanup
  • Image payload sanitization
  • Blank text block cleanup before provider replay
  • Incomplete reasoning only length turn cleanup before provider replay
  • User input provenance tagging (for inter session routed prompts)
  • Empty assistant error turn repair for Bedrock Converse replay

For transcript storage details, see Session management deep dive.


Global rule: runtime context is not user transcript

Runtime or system context can be added to the model prompt for a turn, but it is not content authored by an end user. OpenClaw keeps a separate transcript facing prompt body for Gateway replies, queued followups, ACP, CLI, and embedded OpenClaw runs. Stored visible user turns use that transcript body instead of the runtime enriched prompt.

For legacy sessions that already persisted runtime wrappers, Gateway history surfaces apply a display projection before returning messages to WebChat, TUI, REST, or SSE clients.


Where this runs

All transcript hygiene is centralized in the embedded runner:

  • Policy selection: src/agents/transcript-policy.ts (resolveTranscriptPolicy, keyed on provider, modelApi, and modelId)
  • Sanitization and repair application: sanitizeSessionHistory in src/agents/embedded-agent-runner/replay-history.ts

Separate from transcript hygiene, session files are repaired (if needed) before loading:

  • repairSessionFileIfNeeded in src/agents/session-file-repair.ts
  • Called from src/agents/embedded-agent-runner/run/attempt.ts and src/agents/embedded-agent-runner/compact.ts

Global rule: image sanitization

Image payloads are always sanitized to prevent provider side rejection due to size limits (downscaling and recompressing oversized base64 images). This also helps control image driven token pressure for vision capable models: lower max dimensions reduce token usage, higher dimensions preserve detail.

Implementation:

  • sanitizeSessionMessagesImages in src/agents/embedded-agent-helpers/images.ts
  • sanitizeContentBlocksImages in src/agents/tool-images.ts
  • Max image side is configurable via agents.defaults.imageMaxDimensionPx (default: 1200)
  • Blank text blocks are removed while this pass walks replay content. Assistant turns that become empty are dropped from the replay copy; user and tool result turns that become empty receive a non-empty omitted content placeholder.

Global rule: malformed tool calls

Assistant tool call blocks missing both input and arguments are dropped before model context is built. This prevents provider rejections from partially persisted tool calls (for example, after a rate limit failure).

Implementation:

  • sanitizeToolCallInputs in src/agents/session-transcript-repair.ts
  • Applied in sanitizeSessionHistory (src/agents/embedded-agent-runner/replay-history.ts)

Global rule: tool result pairing

Tool results are paired to tool call occurrences within each assistant turn before provider specific call IDs are rewritten. Provider generated IDs may repeat on later turns, so a result adjacent to a repeated call stays with that occurrence. A displaced result is moved only when exactly one unresolved occurrence can own it; ambiguous extras are dropped and missing occurrences receive synthetic error results.

Implementation: sanitizeToolUseResultPairing in src/agents/session-transcript-repair.ts


Global rule: incomplete or silent reasoning-only turns

Assistant turns are omitted from the in memory replay copy when they contain only thinking or redacted thinking content after either of these events:

  • The provider output limit ends the turn with incomplete reasoning state.
  • Silent reply cleanup removes the turn's only visible NO_REPLY text.

The silent reply cleanup prevents hidden reasoning from merging into a later assistant tool use turn when strict providers rebuild the conversation.

Empty length turns remain unchanged, as do length turns with visible text, tool calls, or unknown content blocks. Silent reply turns with tool calls or unknown content blocks also remain unchanged. Stored transcripts are not rewritten.

Implementation: normalizeAssistantReplayContent in src/agents/embedded-agent-runner/replay-history.ts


Global rule: inter-session input provenance

When an agent sends a prompt into another session via sessions_send (including agent to agent reply and announce steps), OpenClaw persists the created user turn with message.provenance.kind = "inter_session".

OpenClaw also prepends a same turn [Inter-session message] ... isUser=false marker before the routed prompt text so the active model call can distinguish foreign session output from external end user instructions. This marker includes the source session, channel, and tool when available. The transcript still uses role: "user" for provider compatibility, but the visible text and provenance metadata both mark the turn as inter session data.

During context rebuild, OpenClaw applies the same marker to older persisted inter session user turns that only have provenance metadata.


Provider matrix (current behavior)

OpenAI / OpenAI Codex

  • Image sanitization only.
  • Drop orphaned reasoning signatures (standalone reasoning items without a following content block) for OpenAI Responses/Codex transcripts, and drop replayable OpenAI reasoning after a model route switch.
  • Preserve replayable OpenAI Responses reasoning item payloads, including encrypted empty summary items, so manual/WebSocket replay keeps required rs_* state paired with assistant output items.
  • Native ChatGPT Codex Responses follows Codex wire parity by replaying prior Responses reasoning/message/function payloads without prior item IDs while preserving session prompt_cache_key.
  • OpenAI Responses family replay preserves canonical call_*|fc_* same model reasoning pairs, but deterministically normalizes malformed or overlong call_id/function call item ids before pi ai payload conversion.
  • Tool result pairing repair may move real matched outputs and synthesize Codex style aborted outputs for missing tool calls.
  • No turn validation or reordering; no thought signature stripping.

OpenAI compatible Chat Completions

  • Historical assistant thinking and reasoning blocks are stripped before replay so local and proxy style OpenAI compatible servers do not receive prior turn reasoning fields such as reasoning or reasoning_content.
  • Current same turn tool call continuations keep the assistant reasoning block attached to the tool call until the tool result has been replayed.
  • Custom or self hosted model entries with reasoning: true preserve replayed reasoning metadata.
  • Provider owned exceptions can opt out when their wire protocol requires replayed reasoning metadata.

Google (Generative AI / Gemini CLI / Antigravity)

  • Tool call id sanitization: strict alphanumeric.
  • Tool result pairing repair and synthetic tool results.
  • Turn validation (Gemini style turn alternation).
  • Google turn ordering fixup (prepend a tiny user bootstrap if history starts with assistant).
  • Antigravity Claude: normalize thinking signatures; drop unsigned thinking blocks.

Anthropic / Minimax (Anthropic compatible)

  • Tool result pairing repair and synthetic tool results.
  • Turn validation (merge consecutive user turns to satisfy strict alternation).
  • Trailing assistant prefill turns are stripped from outgoing Anthropic Messages payloads when thinking is enabled, including Cloudflare AI Gateway routes.
  • Pre compaction assistant thinking signatures are stripped before provider replay when a session has been compacted. Thinking signatures are cryptographically bound to the conversation prefix at generation time; after compaction the prefix changes (summarized content replaces the original), so replaying the original signatures causes Anthropic to reject the request with "Invalid signature in thinking block". The thinking text is preserved as an unsigned block and then handled by the rule below.
  • Thinking blocks with missing, empty, or blank replay signatures are stripped before provider conversion. If that empties an assistant turn, OpenClaw keeps turn shape with non-empty omitted reasoning text.
  • Older thinking only assistant turns that must be stripped are replaced with non-empty omitted reasoning text so provider adapters do not drop the replay turn.

Amazon Bedrock (Converse API)

  • Empty assistant stream error turns are repaired to a non-empty fallback text block before replay. Bedrock Converse rejects assistant messages with content: [], so persisted assistant turns with stopReason: "error" and empty content are also repaired on disk before load.
  • Assistant stream error turns with only blank text blocks are dropped from the in memory replay copy instead of replaying an invalid blank block.
  • Pre compaction assistant thinking signatures are stripped before Converse replay when a session has been compacted, for the same reason as Anthropic above.
  • Claude thinking blocks with missing, empty, or blank replay signatures are stripped before Converse replay. If that empties an assistant turn, OpenClaw keeps turn shape with non-empty omitted reasoning text.
  • Older thinking only assistant turns that must be stripped are replaced with non-empty omitted reasoning text so the Converse replay keeps strict turn shape.
  • Replay filters OpenClaw delivery mirror and gateway injected assistant turns.
  • Image sanitization applies through the global rule.

Mistral (including model-id based detection)

  • Tool call identifier sanitization: strict9 (alphanumeric, length 9).

OpenRouter Gemini

  • Thought signature cleanup: strip non-base64 thought_signature values (retain base64).

OpenRouter Anthropic

  • Verified OpenRouter OpenAI-compatible Anthropic model payloads have trailing assistant prefill turns removed when reasoning is active, which aligns with how direct Anthropic and Cloudflare Anthropic replays behave.

Everything else

  • Only image sanitization is applied.

Historical behavior (pre-2026.1.22)

Prior to the 2026.1.22 release, OpenClaw applied transcript hygiene at multiple levels:

  • A transcript-sanitize extension executed during every context build and could:
    • Fix tool use and result pairing.
    • Sanitize tool call identifiers (including a non-strict mode that kept _/-).
  • Provider-specific sanitization also ran in the runner, creating redundant effort.
  • Further changes happened outside the provider policy, such as stripping <final> tags from assistant text before saving, discarding empty assistant error turns, and cutting assistant content after tool calls.

This complexity introduced regressions across providers (notably openai-responses call_id|fc_id pairing). The 2026.1.22 cleanup removed the extension, moved all logic into the runner, and made OpenAI no-touch apart from image sanitization.