> ## Documentation Index
> Fetch the complete documentation index at: https://daily-ms-ws-body-url-encode.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Gradium Speech-to-Text

> GradiumSTTService streams real-time STT over Gradium's WebSocket API with multilingual transcription and semantic VAD.

## Overview

`GradiumSTTService` provides real-time speech recognition using Gradium's WebSocket API with support for multilingual transcription, semantic voice activity detection for smart turn-taking, and robust performance in noisy environments.

Transcribes English by default. Set `settings.language` to one of the other supported languages (German, Spanish, French, Portuguese), or to `"any"` to have Gradium detect the language.

By default the pipeline's VAD closes each utterance. With `enable_turn_detection=True` the server's end-pointing signal drives turns instead: the service broadcasts `ProposedUserStartedSpeakingFrame` and `ProposedUserStoppedSpeakingFrame`, and recommends `ExternalUserTurnStrategies` to resolve them.

<CardGroup cols={2}>
  <Card title="Gradium STT API Reference" icon="code" href="https://reference-server.pipecat.ai/en/latest/api/pipecat.services.gradium.stt.html">
    Pipecat's API methods for Gradium STT integration
  </Card>

  <Card title="Example Implementation" icon="play" href="https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-gradium.py">
    Complete example with interruption handling
  </Card>

  <Card title="Turn Detection Example" icon="waveform-lines" href="https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-gradium-turns.py">
    Server-side turn detection example
  </Card>

  <Card title="Gradium Documentation" icon="book" href="https://gradium.ai/api_docs.html#tag/STT">
    Official Gradium STT API documentation
  </Card>

  <Card title="Gradium Platform" icon="microphone" href="https://gradium.ai/">
    Access API keys and speech models
  </Card>
</CardGroup>

## Installation

To use Gradium services, install the required dependency:

```bash theme={null}
uv add "pipecat-ai[gradium]"
```

## Prerequisites

### Gradium Account Setup

Before using Gradium STT services, you need:

1. **Gradium Account**: Sign up at [Gradium](https://gradium.ai/)
2. **API Key**: Generate an API key from your account dashboard
3. **Region Selection**: Choose your preferred region (EU or US)

### Required Environment Variables

* `GRADIUM_API_KEY`: Your Gradium API key for authentication

## Configuration

### GradiumSTTService

<ParamField path="api_key" type="str" required>
  Gradium API key for authentication.
</ParamField>

<ParamField path="api_endpoint_base_url" type="str" default="wss://api.gradium.ai/api/speech/asr">
  WebSocket endpoint URL. Gradium automatically routes traffic to the nearest
  endpoint. Override to pin to a specific region or custom deployment.
</ParamField>

<ParamField path="encoding" type="str" default="pcm">
  Base audio encoding type. One of `"pcm"`, `"wav"`, or `"opus"`. For PCM, the
  sample rate is appended automatically to form the input format (e.g., `"pcm"`
  becomes `"pcm_16000"`). PCM accepts 8000, 16000, and 24000 Hz sample rates.
</ParamField>

<ParamField path="sample_rate" type="int | None" default="None">
  Audio sample rate in Hz. If `None`, uses the pipeline's audio sample rate.
</ParamField>

<ParamField path="params" type="GradiumSTTService.InputParams" default="None" deprecated>
  Configuration parameters for language and delay settings. *Deprecated in
  v0.0.105. Use `settings=GradiumSTTService.Settings(...)` instead.*
</ParamField>

<ParamField path="json_config" type="str" default="None">
  Optional JSON configuration string for additional model settings. Deprecated
  in favor of `params`.
</ParamField>

<ParamField path="enable_turn_detection" type="bool" default="False">
  Whether the server's end-pointing signal decides when user turns start and
  end, instead of the pipeline's VAD. When on, the service proposes turn
  boundaries (`ProposedUserStartedSpeakingFrame` and
  `ProposedUserStoppedSpeakingFrame`) and recommends
  `ExternalUserTurnStrategies` to the user aggregator.
</ParamField>

<ParamField path="settings" type="GradiumSTTService.Settings" default="None">
  Runtime-configurable settings for the STT service. See [Settings](#settings)
  below.
</ParamField>

<ParamField path="ttfs_p99_latency" type="float" default="GRADIUM_TTFS_P99">
  P99 latency from speech end to final transcript in seconds. Override for your
  deployment. See [stt-benchmark](https://github.com/pipecat-ai/stt-benchmark).
</ParamField>

### Settings

Runtime-configurable settings passed via the `settings` constructor argument using `GradiumSTTService.Settings(...)`. These can be updated mid-conversation with `STTUpdateSettingsFrame`. See [Service Settings](/pipecat/fundamentals/service-settings) for details.

The turn detection settings (`eot_horizon_s`, `eot_threshold`, `post_flush_cooldown_frames`) are `None` unless the service is constructed with `enable_turn_detection=True`, which gives them their defaults.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `model` | `str` | `"default"` | STT model identifier. *(Inherited from base STT settings.)* |
| `language` | `Language \| str` | `Language.EN` | Expected language of the audio. *(Inherited from base STT settings.)* Helps ground the model to a specific language and improve transcription quality. Defaults to `Language.EN`; `"any"` asks Gradium to detect the language. |
| `delay_in_frames` | `int` | `12` | Server-side delay in audio frames (80ms each) before text is generated. Higher delays allow more context but increase latency. Allowed values: 7, 8, 10, 12, 14, 16, 20, 24, 36, 48. Default is 12 (960ms). Sent to Gradium API via json\_config. |
| `eot_horizon_s` | `float \| None` | `3.0` | Which end-pointing horizon drives turn decisions; the step entry with the closest `horizon_s` is used. *(Only set when `enable_turn_detection=True`.)* |
| `eot_threshold` | `float \| None` | `0.5` | Inactivity probability on that horizon at or above which an open turn ends, and below which a turn opens once the signal has read inactive. *(Only set when `enable_turn_detection=True`.)* |
| `post_flush_cooldown_frames` | `int \| None` | `8` | Number of step messages to ignore after a flush completes before the next turn may start or end. A flush feeds the model `delay_in_frames` of silence, and the end-pointing signal is unreliable on the frames that follow: it can dip as if speech resumed, then fire again on the first real frames. *(Only set when `enable_turn_detection=True`.)* |

## Usage

### Basic Setup

```python theme={null}
from pipecat.services.gradium.stt import GradiumSTTService

stt = GradiumSTTService(
    api_key=os.getenv("GRADIUM_API_KEY"),
)
```

### With Language and Delay Configuration

```python theme={null}
from pipecat.services.gradium.stt import GradiumSTTService
from pipecat.transcriptions.language import Language

stt = GradiumSTTService(
    api_key=os.getenv("GRADIUM_API_KEY"),
    settings=GradiumSTTService.Settings(
        language=Language.EN,
        delay_in_frames=8,
    ),
)
```

### With Server-Side Turn Detection

```python theme={null}
from pipecat.services.gradium.stt import GradiumSTTService

stt = GradiumSTTService(
    api_key=os.getenv("GRADIUM_API_KEY"),
    enable_turn_detection=True,
    settings=GradiumSTTService.Settings(
        eot_horizon_s=3.0,
        eot_threshold=0.5,
    ),
)
```

## Notes

* **Supported languages**: German, English, Spanish, French, and Portuguese.
* **Audio format**: Configurable via `encoding` and `sample_rate` parameters. Defaults to PCM with the pipeline's sample rate. Supported PCM rates: 8000, 16000, and 24000 Hz. Audio is sent in 80ms chunks.
* **Turn detection**: With `enable_turn_detection=True`, the server's end-pointing signal decides when user turns start and end. The service broadcasts `ProposedUserStartedSpeakingFrame` when a turn opens and `ProposedUserStoppedSpeakingFrame` when it closes, then flushes the server. The final `TranscriptionFrame` follows once the flush is acknowledged. Local VAD frames are ignored. The service recommends `ExternalUserTurnStrategies`, which resolve the proposals into user turn frames and hold the turn open until the transcript arrives.

<Tip>
  The `InputParams` / `params=` pattern is deprecated as of v0.0.105. Use
  `Settings` / `settings=` instead. See the [Service Settings
  guide](/pipecat/fundamentals/service-settings) for migration details.
</Tip>

## Event Handlers

Gradium STT supports the standard [service connection events](/api-reference/server/events/service-events):

| Event | Description |
| - | - |
| `on_connected` | Connected to Gradium WebSocket |
| `on_disconnected` | Disconnected from Gradium WebSocket |
| `on_turn_start` | The end-pointing signal opened a turn *(turn detection only)* |
| `on_turn_end` | The end-pointing signal closed the turn *(turn detection only)* |

```python theme={null}
@stt.event_handler("on_connected")
async def on_connected(service):
    print("Connected to Gradium")

@stt.event_handler("on_turn_start")
async def on_turn_start(service):
    print("Turn started")

@stt.event_handler("on_turn_end")
async def on_turn_end(service):
    print("Turn ended")
```
