Deepgram Speech-to-Text API for Audio Transcription and Voice Calls

Learn how to configure Deepgram for batch transcription of inbound voice notes and streaming STT for voice calls in OpenClaw. This page covers API key setup, provider enabling, and usage.

Read this when

  • You want Deepgram speech-to-text for audio attachments
  • You want Deepgram streaming transcription for Voice Call
  • You need a quick Deepgram config example

Deepgram is a speech-to-text API. OpenClaw uses it to transcribe inbound audio and voice notes through tools.media.audio and to perform streaming STT for Voice Calls via plugins.entries.voice-call.config.streaming.

Batch transcription uploads the full audio file to Deepgram and inserts the transcript into the reply pipeline using the {{Transcript}} and [Audio] blocks. Voice Call streaming sends live G.711 u-law frames to Deepgram's WebSocket listen endpoint and outputs partial or final transcripts as Deepgram delivers them.

DetailValue
Websitedeepgram.com
Docsdevelopers.deepgram.com
AuthDEEPGRAM_API_KEY
Default modelnova-3

Getting started

Set your API key

DEEPGRAM_API_KEY=dg_...

Enable the audio provider

{
  tools: {
    media: {
      audio: {
        enabled: true,
        models: [{ provider: "deepgram", model: "nova-3" }],
      },
    },
  },
}

Send a voice note

Send an audio message through any connected channel. OpenClaw transcribes it using Deepgram and inserts the transcript into the reply pipeline.

Configuration options

OptionPathDescription
modeltools.media.models[].modelDeepgram model id (default: nova-3)
languagetools.media.models[].languageLanguage hint (optional)

providerOptions.deepgram appends extra query parameters directly to the Deepgram /listen request, so any parameter supported by Deepgram works (for example detect_language, punctuate, smart_format):

With language hint

{
  tools: {
    media: {
      audio: {
        enabled: true,
        models: [{ provider: "deepgram", model: "nova-3", language: "en" }],
      },
    },
  },
}

With Deepgram options

{
  tools: {
    media: {
      audio: {
        enabled: true,
        providerOptions: {
          deepgram: {
            detect_language: true,
            punctuate: true,
            smart_format: true,
          },
        },
        models: [{ provider: "deepgram", model: "nova-3" }],
      },
    },
  },
}

Voice Call streaming STT

The bundled deepgram plugin also registers a realtime transcription provider for the Voice Call plugin.

SettingConfig pathDefault
API keyplugins.entries.voice-call.config.streaming.providers.deepgram.apiKeyFalls back to DEEPGRAM_API_KEY
Base URL...deepgram.baseUrlDEEPGRAM_BASE_URL or Deepgram's public API
Model...deepgram.modelnova-3
Language...deepgram.language(unset)
Encoding...deepgram.encodingmulaw
Sample rate...deepgram.sampleRate8000
Endpointing...deepgram.endpointingMs800
Interim results...deepgram.interimResultstrue
{
  plugins: {
    entries: {
      "voice-call": {
        config: {
          streaming: {
            enabled: true,
            provider: "deepgram",
            providers: {
              deepgram: {
                apiKey: "${DEEPGRAM_API_KEY}",
                model: "nova-3",
                endpointingMs: 800,
                language: "en-US",
              },
            },
          },
        },
      },
    },
  },
}

For a Deepgram custom endpoint, set baseUrl to the endpoint root, including any base path but not /listen. Realtime endpoints accept http://, https://, ws://, and wss://. HTTP maps to WS, HTTPS maps to WSS, and explicit WebSocket schemes stay unchanged. Malformed URLs and other schemes fail during session setup.

Note

Voice Call receives telephony audio as 8 kHz G.711 u-law. The Deepgram streaming provider defaults to encoding: "mulaw" and sampleRate: 8000, so Twilio media frames can be forwarded directly.

Notes

Authentication

Authentication follows the standard provider auth order. DEEPGRAM_API_KEY is the simplest path.

Proxy and custom endpoints

Override endpoints or headers on the Deepgram tools.media.models[] entry when using a proxy.

Output behavior

Output follows the same audio rules as other providers (size caps, timeouts, transcript injection).

  • Media tools, Audio, image, and video processing pipeline overview.

  • Configuration, Full config reference including media tool settings.

  • Troubleshooting, Common issues and debugging steps.

  • FAQ, Frequently asked questions about OpenClaw setup.