Inworld Streaming TTS Provider for OpenClaw Replies

This page covers the Inworld streaming text-to-speech provider for OpenClaw, including setup, configuration, and output formats. Developers integrating TTS into OpenClaw replies will find this guide essential.

Read this when

  • You want Inworld speech synthesis for outbound replies
  • You need PCM telephony or OGG_OPUS voice-note output from Inworld

Inworld functions as a streaming text-to-speech (TTS) provider. Within OpenClaw, it generates outbound reply audio, producing MP3 by default, OGG_OPUS for voice notes, and raw PCM audio for telephony channels like Voice Call.

OpenClaw sends requests to Inworld's streaming TTS endpoint, merges the returned base64 audio chunks into one buffer, then passes the result into the standard reply audio pipeline.

PropertyValue
Provider idinworld
Pluginofficial external package (@openclaw/inworld-speech)
ContractspeechProviders (TTS only)
Auth env varINWORLD_API_KEY (HTTP Basic, Base64 dashboard credential)
Base URLhttps://api.inworld.ai
Default voiceSarah
Default modelinworld-tts-1.5-max
OutputMP3 (default), OGG_OPUS (voice notes), PCM 22050 Hz (telephony)
Websiteinworld.ai
Docsdocs.inworld.ai/tts/tts

Install plugin

openclaw plugins install @openclaw/inworld-speech
openclaw gateway restart

Getting started

Set your API key

Retrieve the credential from your Inworld dashboard under Workspace > API Keys and assign it as an environment variable. The value is transmitted exactly as the HTTP Basic credential, so avoid Base64 re-encoding or converting it into a bearer token.

INWORLD_API_KEY=<base64-credential-from-dashboard>

Select Inworld in tts

{
  tts: {
    auto: "always",
    provider: "inworld",
    providers: {
      inworld: {
        voiceId: "Sarah",
        modelId: "inworld-tts-1.5-max",
      },
    },
  },
}

Send a message

Send a reply through any connected channel. OpenClaw generates the audio using Inworld and delivers it as MP3, or OGG_OPUS when the channel expects a voice note.

Configuration options

OptionPathDescription
apiKeytts.providers.inworld.apiKeyBase64 dashboard credential. Falls back to INWORLD_API_KEY.
baseUrltts.providers.inworld.baseUrlOverride Inworld API base URL (default https://api.inworld.ai).
voiceIdtts.providers.inworld.voiceIdVoice identifier (default Sarah). Legacy alias: speakerVoiceId.
modelIdtts.providers.inworld.modelIdTTS model id (default inworld-tts-1.5-max).
temperaturetts.providers.inworld.temperatureSampling temperature, 0 (exclusive) to 2 (optional).

Notes

Authentication

Inworld relies on HTTP Basic authentication using a single Base64-encoded credential string. Copy it directly from the Inworld dashboard. The provider sends it as Authorization: Basic <apiKey> without additional encoding, so do not Base64-encode it yourself and avoid passing a bearer-style token. The same reminder appears in the TTS auth notes.

Models

Supported model identifiers: inworld-tts-1.5-max (default), inworld-tts-1.5-mini, inworld-tts-1-max, inworld-tts-1.

Audio outputs

Replies default to MP3. When the channel target is voice-note, OpenClaw requests OGG_OPUS from Inworld so the audio plays as a native voice bubble. Telephony synthesis uses raw PCM at 22050 Hz to supply the telephony bridge.

Custom endpoints

Override the API host using tts.providers.inworld.baseUrl. Trailing slashes are removed before requests are dispatched.

498 words · updated Jul 27, 2026