> ## Documentation Index
> Fetch the complete documentation index at: https://daily-ms-ws-body-url-encode.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Speechmatics Speech-to-Text

> SpeechmaticsSTTService delivers real-time STT with partial and final results, speaker diarization, and turn detection via Agent STT API.

## Overview

`SpeechmaticsSTTService` enables real-time speech transcription using Speechmatics' Agent STT WebSocket API with partial and final results, speaker diarization, and server-side or client-side turn detection.

<Note>
  Speechmatics provides its own user turn start and end detection. When using
  `TurnDetectionMode.VAD`, the service runs its own VAD and closes turns itself,
  automatically requesting `ExternalUserTurnStrategies` at start. The service
  proposes each turn boundary and those strategies resolve it, so
  `should_interrupt` is the control for barge-in. See [User Turn
  Strategies](/api-reference/server/utilities/turn-management/user-turn-strategies)
  for more details. A VAD in the transport (such as `SileroVADAnalyzer`) is
  optional when Speechmatics drives turn detection; include it if you want
  useful STT metrics.
</Note>

<CardGroup cols={2}>
  <Card title="Speechmatics Agent STT Reference" icon="code" href="https://docs.speechmatics.com/api-ref/agent-stt-websocket">
    The Agent STT WebSocket API this service speaks
  </Card>

  <Card title="Example Implementation" icon="play" href="https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-speechmatics.py">
    Complete example with interruption handling
  </Card>

  <Card title="Speechmatics Documentation" icon="book" href="https://docs.speechmatics.com/integrations-and-sdks/pipecat/stt">
    Speechmatics' own guide to this integration
  </Card>

  <Card title="Speaker Diarization Guide" icon="microphone" href="https://docs.speechmatics.com/speech-to-text/features/diarization#speaker-diarization">
    Learn about separating different speakers in audio
  </Card>
</CardGroup>

## Installation

To use Speechmatics services, install the required dependencies:

```bash theme={null}
uv add "pipecat-ai[speechmatics]~=1.10.0"
```

## Prerequisites

### Speechmatics Account Setup

Before using Speechmatics STT services, you need:

1. **Speechmatics Account**: Sign up at [Speechmatics](https://www.speechmatics.com/)
2. **API Key**: Generate an API key from your account dashboard
3. **Feature Selection**: Configure transcription features like speaker diarization

### Select Endpoint

Speechmatics Agent STT supports the following endpoints (defaults to `EU2`):

| Region | Environment | STT Endpoint |
| - | - | - |
| EU | EU2 (Default) | `wss://eu2.rt.speechmatics.com/v2/agent` |
| All regions | Global | `wss://global.rt.speechmatics.com/v2/agent` |
| EU | EU1 | `wss://eu.rt.speechmatics.com/v2/agent` |
| US | US1 | `wss://us.rt.speechmatics.com/v2/agent` |
| AU | AU1 | `wss://au.rt.speechmatics.com/v2/agent` |

`global.rt.speechmatics.com` routes each connection to the nearest region for lowest latency. It may route to any region, so use a regional endpoint if you have data residency requirements. For the full list, see [Supported endpoints](https://docs.speechmatics.com/get-started/authentication#supported-endpoints).

### Required Environment Variables

* `SPEECHMATICS_API_KEY`: Your Speechmatics API key for authentication
* `SPEECHMATICS_RT_URL`: Speechmatics endpoint URL (optional, defaults to EU2)

## Configuration

### SpeechmaticsSTTService

<ParamField path="api_key" type="str" default="None">
  Speechmatics API key. Falls back to the `SPEECHMATICS_API_KEY` environment
  variable.
</ParamField>

<ParamField path="base_url" type="str" default="None">
  Base URL for the Speechmatics API. Falls back to `SPEECHMATICS_RT_URL`
  environment variable, then defaults to
  `wss://eu2.rt.speechmatics.com/v2/agent`.
</ParamField>

<ParamField path="sample_rate" type="int" default="None">
  Audio sample rate in Hz. When `None`, uses the pipeline's configured sample
  rate.
</ParamField>

<ParamField path="encoding" type="AudioEncoding" default="AudioEncoding.PCM_S16LE">
  Audio encoding format. Init-only -- not part of runtime-updatable settings.
</ParamField>

<ParamField path="params" type="SpeechmaticsSTTService.InputParams" default="None" deprecated>
  Additional configuration parameters. *Deprecated in v0.0.105. Use
  `settings=SpeechmaticsSTTService.Settings(...)` instead.*
</ParamField>

<ParamField path="settings" type="SpeechmaticsSTTService.Settings" default="None">
  Runtime-configurable settings for the STT service. See [Settings](#settings)
  below.
</ParamField>

<ParamField path="should_interrupt" type="bool" default="True">
  Whether to interrupt bot output when Speechmatics detects user speech. Only
  applies when `turn_detection_mode` is set to `VAD`. See [User Turn
  Strategies](/api-reference/server/utilities/turn-management/user-turn-strategies#externaluserturnstrategies)
  if you pass your own `user_turn_strategies`.
</ParamField>

<ParamField path="ttfs_p99_latency" type="float" default="0.74">
  P99 latency from speech end to final transcript in seconds. Override for your
  deployment. See [stt-benchmark](https://github.com/pipecat-ai/stt-benchmark).
</ParamField>

### Settings

Runtime-configurable settings passed via the `settings` constructor argument using `SpeechmaticsSTTService.Settings(...)`. These can be updated mid-conversation with `STTUpdateSettingsFrame`. See [Service Settings](/pipecat/fundamentals/service-settings) for details.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `model` | `str` | `"linden-1"` | Transcription model to use. `linden-1` is the only model Agent STT accepts. *(Inherited from base STT settings.)* Supersedes `operating_point`. |
| `language` | `Language \| str` | `Language.EN` | Language code for transcription. *(Inherited from base STT settings.)* |
| `domain` | `str` | `None` | Domain for Speechmatics API (e.g. for bilingual transcription). |
| `turn_detection_mode` | `TurnDetectionMode` | `EXTERNAL` | How turns are closed. `EXTERNAL` has the caller drive turns via `finalize()` (e.g. Pipecat's own VAD); `VAD` lets the STT service run its own VAD and close turns itself. |
| `speaker_active_format` | `str` | `"{text}"` | Formatter for speaker output. Available attributes: `{speaker_id}`, `{text}`, `{ts}`, `{lang}`. Example: `"@{speaker_id}: {text}"`. |
| `known_speakers` | `list[SpeakerIdentifier]` | `[]` | Known speaker labels and identifiers for speaker attribution. |
| `additional_vocab` | `list[AdditionalVocabEntry]` | `[]` | Additional vocabulary to boost recognition of specific words. |
| `operating_point` | `Model \| str` | `None` | *Deprecated since 1.10.0.* Resolves to `model`, warning when used. If both are given they must name the same value, otherwise a `ValueError` is raised. Removed in 2.0.0. |
| `enable_partials` | `bool` | `None` | Emit interim segments as `AddPartialSegment`, updating until the final segment arrives. Agent STT enables partials by default. Set to false to disable partials. |
| `punctuation_overrides` | `dict` | `None` | Custom punctuation overrides for the STT engine. |
| `enable_diarization` | `bool` | `None` | Enable speaker diarization to attribute words to unique speakers. |
| `speaker_sensitivity` | `float` | `None` | Diarization sensitivity. Higher values help distinguish similar voices. |
| `max_speakers` | `int` | `None` | Maximum number of speakers to detect. Only use when the speaker count is known. |
| `prefer_current_speaker` | `bool` | `None` | Give extra weight to grouping nearby words as the same speaker. |

## Turn Detection

The Speechmatics Agent STT service supports two turn detection modes:

### External mode (Default)

The default `TurnDetectionMode.EXTERNAL` hands endpointing to the caller. The service does not endpoint on its own; instead, the caller drives turns by calling `finalize()`, typically from Pipecat's own VAD. Use this mode when you want Pipecat's VAD to control turn boundaries.

```python theme={null}
transport_params = TransportParams(
    audio_in_enabled=True,
    audio_out_enabled=True,
    vad_analyzer=SileroVADAnalyzer(),  # Use Pipecat's VAD
)

...

stt = SpeechmaticsSTTService(
    api_key=os.getenv("SPEECHMATICS_API_KEY"),
    settings=SpeechmaticsSTTService.Settings(
        language=Language.EN,
        speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",
    ),
)
```

### VAD mode

`TurnDetectionMode.VAD` lets the Speechmatics service run its own VAD and close turns itself. The service emits `StartOfTurn` and `EndOfTurn` events, proposes turn boundaries to Pipecat's turn strategies, and handles endpointing server-side. In this mode, you do not need a VAD in the transport — the service manages turn detection.

```python theme={null}
stt = SpeechmaticsSTTService(
    api_key=os.getenv("SPEECHMATICS_API_KEY"),
    settings=SpeechmaticsSTTService.Settings(
        language=Language.EN,
        turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.VAD,
        speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",
    ),
)
```

<Note>
  When using `VAD` mode, remove any `vad_analyzer` or `turn_analyzer` from your
  transport configuration to avoid conflicts — Speechmatics handles turn
  detection itself.
</Note>

## Speaker Diarization

Speechmatics STT supports speaker diarization, which separates out different speakers in the audio. The identity of each speaker is returned in the TranscriptionFrame objects in the `user_id` attribute.

If `speaker_active_format` is provided, then the text output for the TranscriptionFrame will be formatted to this specification. Your system context can then be updated to include information about this format to understand which speaker spoke which words.

Examples:

* `<{speaker_id}>{text}</{speaker_id}>` → `<S1>Good morning.</S1>`.
* `@{speaker_id}: {text}` → `@S1: Good morning.`.

### Available attributes

| Attribute | Description | Example |
| - | - | - |
| `speaker_id` | The label of the speaker | `S1` |
| `text` | The transcribed text | `Good morning.` |
| `ts` | Start time of the segment, in seconds | `12.34` |
| `lang` | The language of the transcription | `en` |

## Language Support

<Note>
  Refer to the [Speechmatics
  docs](https://docs.speechmatics.com/speech-to-text/languages) for more
  information on supported languages.
</Note>

Speechmatics STT supports the following languages and regional variants.

Setting a language can be done using the `language` parameter when creating the STT object. The exception to this is English / Mandarin which has the code `cmn_en`.

| Language Code | Description | Locales |
| - | - | - |
| `Language.AR` | Arabic | - |
| `Language.BA` | Bashkir | - |
| `Language.EU` | Basque | - |
| `Language.BE` | Belarusian | - |
| `Language.BG` | Bulgarian | - |
| `Language.BN` | Bengali | - |
| `Language.YUE` | Cantonese | - |
| `Language.CA` | Catalan | - |
| `Language.HR` | Croatian | - |
| `Language.CS` | Czech | - |
| `Language.DA` | Danish | - |
| `Language.NL` | Dutch | - |
| `Language.EN` | English | `en-US`, `en-GB`, `en-AU` |
| `Language.EO` | Esperanto | - |
| `Language.ET` | Estonian | - |
| `Language.FA` | Persian | - |
| `Language.FI` | Finnish | - |
| `Language.FIL` | Filipino | - |
| `Language.FR` | French | - |
| `Language.GL` | Galician | - |
| `Language.DE` | German | - |
| `Language.EL` | Greek | - |
| `Language.HE` | Hebrew | - |
| `Language.HI` | Hindi | - |
| `Language.HU` | Hungarian | - |
| `Language.IA` | Interlingua | - |
| `Language.IT` | Italian | - |
| `Language.ID` | Indonesian | - |
| `Language.GA` | Irish | - |
| `Language.JA` | Japanese | - |
| `Language.KO` | Korean | - |
| `Language.LV` | Latvian | - |
| `Language.LT` | Lithuanian | - |
| `Language.MS` | Malay | - |
| `Language.MT` | Maltese | - |
| `Language.CMN` | Mandarin | `cmn-Hans`, `cmn-Hant` |
| `Language.MR` | Marathi | - |
| `Language.MN` | Mongolian | - |
| `Language.NO` | Norwegian | - |
| `Language.PL` | Polish | - |
| `Language.PT` | Portuguese | - |
| `Language.RO` | Romanian | - |
| `Language.RU` | Russian | - |
| `Language.SK` | Slovakian | - |
| `Language.SL` | Slovenian | - |
| `Language.ES` | Spanish | - |
| `Language.SV` | Swedish | - |
| `Language.SW` | Swahili | - |
| `Language.TA` | Tamil | - |
| `Language.TL` | Tagalog | - |
| `Language.TH` | Thai | - |
| `Language.TR` | Turkish | - |
| `Language.UG` | Uyghur | - |
| `Language.UK` | Ukrainian | - |
| `Language.UR` | Urdu | - |
| `Language.VI` | Vietnamese | - |
| `Language.CY` | Welsh | - |

For bilingual transcription, use the `language` and `domain` parameters as follows:

| Language Code | Description | Domain Options |
| - | - | - |
| `cmn_en` | English / Mandarin | - |
| `en_ms` | English / Malay | - |
| `Language.ES` | English / Spanish | `bilingual-en` |
| `en_ta` | English / Tamil | - |

## Usage Examples

Examples are included in the Pipecat project:

* Using Speechmatics STT service -> [07a-interruptible-speechmatics.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-speechmatics.py)
* Using Speechmatics STT service with VAD -> [07a-interruptible-speechmatics-vad.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/voice/voice-speechmatics-vad.py)
* Transcribing with Speechmatics STT -> [13h-speechmatics-transcription.py](https://github.com/pipecat-ai/pipecat/blob/main/examples/transcription/transcription-speechmatics.py)

Sample projects:

* Guess Who -> [Guess Who](https://github.com/sam-s10s/pipecat-guess-who)
* Guess Who Board Game -> [Guess Who](https://github.com/sam-s10s/pipecat-guess-who-irl)

### Basic Configuration

Initialize the `SpeechmaticsSTTService` and use it in a pipeline:

```python theme={null}
from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

# Configure service
stt = SpeechmaticsSTTService(
    api_key="your-api-key",
    settings=SpeechmaticsSTTService.Settings(
        language=Language.FR,
    )
)

# Use in pipeline
pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    transport.output(),
    context_aggregator.assistant()
])
```

### With Diarization

Enable diarization to attribute transcribed words to unique speakers.

Initialize the `SpeechmaticsSTTService` and use it in a pipeline:

```python theme={null}
from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

# Configure service
stt = SpeechmaticsSTTService(
    api_key="your-api-key",
    settings=SpeechmaticsSTTService.Settings(
        language=Language.EN,
        turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.VAD,
        enable_diarization=True,
        speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",
    )
)

# Use in pipeline
pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    transport.output(),
    context_aggregator.assistant()
])
```

## Additional Notes

* **Connection Management**: Automatically handles WebSocket connections and reconnections with exponential backoff. Transient failures trigger automatic reconnection while audio is buffered; permanent errors (auth rejection, bad configuration) mark the service unusable.
* **Sample Rate**: The default sample rate is `16000` Hz in `pcm_s16le` format
* **VAD Integration**: Server-side VAD and turn detection are available in `VAD` mode

<Tip>
  The `InputParams` / `params=` pattern is deprecated as of v0.0.105. Use
  `Settings` / `settings=` instead. See the [Service Settings
  guide](/pipecat/fundamentals/service-settings) for migration details.
</Tip>

## Event Handlers

In addition to the standard [service connection events](/api-reference/server/events/service-events) (`on_connected`, `on_disconnected`, `on_connection_error`), Speechmatics provides:

| Event | Description |
| - | - |
| `on_speakers_result` | Speaker identification result received |

```python theme={null}
@stt.event_handler("on_speakers_result")
async def on_speakers_result(service, message):
    print(f"Speaker result: {message}")
```
