Media Overview: Image, Video, Music, and Speech Tools

This page covers OpenClaw's media capabilities, including image, video, music, and speech tools. It's for users who want to understand how media generation and processing work within chat sessions.

Read this when

  • Looking for an overview of OpenClaw's media capabilities
  • Deciding which media provider to configure
  • Understanding how async media generation works

OpenClaw can produce images, videos, and music, process incoming media such as pictures, sound files, and clips, and respond vocally using text-to-speech. Every media feature depends on tools: the agent picks when to invoke them based on the ongoing chat, and a tool is only exposed once at least one supporting provider is set up.

For live speech, the Talk session contract is used rather than the single-shot media tool flow. Talk offers three modes: provider-native realtime, local or streaming stt-tts, and transcription for observe-only speech capture. These modes align with telephony, meetings, browser realtime, and native push-to-talk clients in terms of provider catalogs, event envelopes, and cancellation behavior.

Capabilities

  • Image generation, Build or modify images from text prompts or reference pictures through image_generate. Runs asynchronously in chat sessions, executing in the background and publishing the output once complete.

  • Video generation, Supports text-to-video, image-to-video, and video-to-video via video_generate. Also asynchronous, finishing in the background and delivering the result when done.

  • Music generation, Produce music or audio tracks using music_generate. Asynchronous in chat sessions, relying on the shared media-generation task lifecycle.

  • Text-to-speech, Turn outbound replies into spoken audio with the tts tool plus tts configuration. This one is synchronous.

  • Media understanding, Summarize incoming images, audio, and video using vision-capable model providers and dedicated media-understanding plugins.

  • Speech-to-text, Convert inbound voice messages to text via batch STT or Voice Call streaming STT providers.

  • Media playback, Play assistant audio and video inline in the Control UI and native apps, with managed access and portable playback renditions.

Provider capability matrix

Note

This table covers the dedicated media-generation, TTS, and STT plugins. Many chat-model providers (Anthropic, Google, OpenAI, and others) also understand inbound media through their reply model; see the full provider list in Media understanding.

ProviderImageVideoMusicTTSSTTRealtime voiceMedia understanding
Alibaba
Azure Speech
BytePlus
ComfyUI
Deepgram
DeepInfra
ElevenLabs
fal
Google
Gradium
Inworld
LiteLLM
Local CLI
Microsoft
Microsoft Foundry
MiniMax
Mistral
OpenAI
OpenRouter
PixVerse
Qwen
Runway
SenseAudio
Together
Volcengine
Vydra
xAI
Xiaomi MiMo

Note

Realtime voice here means provider-native bidirectional realtime (Talk realtime mode, e.g. Gemini Live or the OpenAI Realtime API), only Google and OpenAI register it today. Deepgram, ElevenLabs, Mistral, OpenAI, and xAI separately register Voice Call streaming STT (one-way audio-to-text); see Speech-to-text and Voice Call below. xAI Realtime voice is an upstream capability but is not registered in OpenClaw until the shared realtime-voice contract can represent it.

Async vs synchronous

CapabilityModeWhy
ImageAsynchronousProvider processing can outlive a chat turn; generated attachments use the shared completion path.
Text-to-speechSynchronousProvider responses return in seconds; attached to the reply audio.
VideoAsynchronousProvider processing takes 30 s to several minutes; slow queues can run up to the configured timeout.
MusicAsynchronousSame provider-processing characteristic as video.

For async tools, OpenClaw submits the request to the provider, returns a task id immediately, and tracks the job in the task ledger. The agent continues responding to other messages while the job runs. When the provider finishes, OpenClaw wakes the agent with the generated media paths so it can tell the user through the session's normal visible-reply mode: automatic final reply delivery when configured, or message(action="send") when the session requires the message tool. If the requester session is inactive or its active wake fails, and some generated media is still missing from the completion reply, OpenClaw sends an idempotent direct fallback with only the missing media. Media already delivered by the completion reply is not posted again.

Speech-to-text and Voice Call

Deepgram, DeepInfra, ElevenLabs, Google, Groq, Mistral, OpenAI, OpenRouter, SenseAudio, and xAI can all transcribe inbound audio through the batch tools.media.audio path when configured. Channel plugins that preflight a voice note for mention gating or command parsing mark the transcribed attachment on the inbound context, so the shared media-understanding pass reuses that transcript instead of making a second STT call for the same audio.

Deepgram, ElevenLabs, Mistral, OpenAI, and xAI also register Voice Call streaming STT providers, so live phone audio can be forwarded to the selected vendor without waiting for a completed recording.

For live user conversations, prefer Talk mode. Batch audio attachments stay on the media path; browser realtime, native push-to-talk, telephony, and meeting audio should use Talk events and the session-scoped catalogs returned by the Gateway.

Provider mappings (how vendors split across surfaces)

Google

Image, video, music, batch TTS, batch STT, backend realtime voice, and media-understanding surfaces.

OpenAI

Image, video, batch TTS, batch STT, Voice Call streaming STT, backend realtime voice, and memory-embedding surfaces.

DeepInfra

Chat/model routing, image generation/editing, text-to-video, batch TTS, batch STT, image media understanding, and memory-embedding surfaces. DeepInfra also exposes reranking, classification, object-detection, and other native model types; OpenClaw has no provider contract for those categories yet, so this plugin does not register them.

xAI

Image, video, search, code-execution, batch TTS, batch STT, and Voice Call streaming STT. xAI Realtime voice is an upstream capability but is not registered in OpenClaw until the shared realtime-voice contract can represent it.

1,249 words · updated Aug 1, 2026