Azure Speech Provider for Text-to-Speech in OpenClaw

Learn how to configure and use Azure AI Speech text-to-speech with OpenClaw for MP3, Ogg/Opus, and mulaw output. This guide covers authentication, supported formats, and REST API integration.

Read this when

  • You want Azure Speech synthesis for outbound replies
  • You need native Ogg Opus voice-note output from Azure Speech

Azure Speech is a bundled provider for Azure AI Speech text-to-speech. OpenClaw communicates directly with the Azure Speech REST API using SSML, producing MP3 for standard replies, native Ogg/Opus for voice notes, and 8 kHz mulaw for telephony channels like Voice Call. The request includes the provider-owned output format via the X-Microsoft-OutputFormat header.

DetailValue
Provider IDazure-speech (alias: azure)
WebsiteAzure AI Speech
DocsSpeech REST text-to-speech
AuthAZURE_SPEECH_KEY plus AZURE_SPEECH_REGION
Default voiceen-US-JennyNeural
Default file outputaudio-24khz-48kbitrate-mono-mp3
Default voice-note fileogg-24khz-16bit-mono-opus

Getting started

Create an Azure Speech resource

In the Azure portal, set up a Speech resource. Copy KEY 1 from Resource Management > Keys and Endpoint, and copy the resource location, for example eastus.

AZURE_SPEECH_KEY=<speech-resource-key>
AZURE_SPEECH_REGION=eastus

Select Azure Speech in tts

{
  tts: {
    auto: "always",
    provider: "azure-speech",
    providers: {
      "azure-speech": {
        voice: "en-US-JennyNeural",
        lang: "en-US",
      },
    },
  },
}

Send a message

Send a reply over any connected channel. OpenClaw generates the audio using Azure Speech and delivers MP3 for standard audio, or Ogg/Opus when the channel expects a voice note.

Configuration options

All options are located under tts.providers["azure-speech"].

OptionDescription
apiKeyAzure Speech resource key. Falls back to AZURE_SPEECH_KEY, AZURE_SPEECH_API_KEY, or SPEECH_KEY.
regionAzure Speech resource region. Falls back to AZURE_SPEECH_REGION or SPEECH_REGION.
endpointOptional Azure Speech endpoint override. Falls back to trusted AZURE_SPEECH_ENDPOINT.
baseUrlOptional Azure Speech base URL override.
voiceAzure voice ShortName (default en-US-JennyNeural). Legacy alias: voiceId.
langSSML language code (default en-US).
outputFormatAudio-file output format (default audio-24khz-48kbitrate-mono-mp3).
voiceNoteOutputFormatVoice-note output format (default ogg-24khz-16bit-mono-opus).
timeoutMsRequest timeout override in milliseconds. Falls back to the global tts.timeoutMs.

The provider is considered configured once apiKey is set along with one of region, endpoint, or baseUrl. Environment variables are only checked as a fallback for config keys that are not set. Workspace .env files cannot set AZURE_SPEECH_ENDPOINT; use the process environment, global runtime dotenv, or explicit config for endpoint routing.

Notes

Authentication

Azure Speech requires a Speech resource key, not an Azure OpenAI key. The key is sent as Ocp-Apim-Subscription-Key; OpenClaw derives https://<region>.tts.speech.microsoft.com from region unless you supply endpoint or baseUrl.

Voice names

Use the Azure Speech voice ShortName value, for example en-US-JennyNeural. The bundled provider can list voices through the same Speech resource and excludes voices marked as deprecated, retired, or disabled.

Audio outputs

Azure accepts output formats like audio-24khz-48kbitrate-mono-mp3, ogg-24khz-16bit-mono-opus, and riff-24khz-16bit-mono-pcm. OpenClaw requests Ogg/Opus for voice-note targets so channels can send native voice bubbles without an extra MP3 conversion, and forces raw-8khz-8bit-mono-mulaw for telephony targets.

Alias

azure is accepted as a provider alias for existing config, but new config should use azure-speech to avoid confusion with Azure OpenAI model providers.