Skip to main content

Overview

VoicemailDetector decides whether an outbound call reached a person or went to voicemail. It listens to what the other side says and asks a classifier. Until it has a verdict, its TTSGate holds the bot’s speech back, so a voicemail greeting is never talked over and a person hears the bot’s reply as soon as the verdict is in. For a guide to setting it up, see Voicemail Detection.

Voicemail Detection Guide

Set up voicemail detection and leave a message

Example Implementation

A complete outbound bot with voicemail detection

Configuration

BaseClassifier
required
What decides between a person and a voicemail, such as a JevClassifier or an LLMClassifier. The detector owns its lifecycle: it sets the classifier up and cleans it up with the pipeline.
float
default:"1.0"
Seconds of silence after the caller stops speaking before the latest answer decides. A greeting resumes after its pauses while a person stops and waits, so a verdict on a fragment such as “hi, this is Sam” is not acted on until the caller has really stopped.
float
default:"2.0"
Seconds of silence after a voicemail verdict before on_voicemail_detected fires, so the message is left after the greeting ends and the recording starts. Speech, such as the rest of the greeting, restarts the wait.
LLMService | None
default:"None"
deprecated
Deprecated since 1.12.0, removed in 2.0.0. An LLM service for the classification. It is wrapped in an LLMClassifier. Pass classifier=LLMClassifier(llm=llm) instead.
str | None
default:"None"
deprecated
Deprecated since 1.12.0, removed in 2.0.0. A system prompt for the llm, placed in front of the LLMClassifier’s own instructions. Pass an LLMClassifier with your own instructions instead.

Methods

detector

The processor to place after the STT service and before the user context aggregator. It is the detector itself.

gate

The TTSGate to place right after the TTS service. While the decision is pending it buffers TTS frames (TTSStartedFrame, TTSTextFrame, TTSAudioRawFrame, TTSStoppedFrame) and passes everything else. A conversation verdict releases the buffered frames in order. A voicemail verdict discards them, since they were meant for a person.

How the Decision Is Made

  1. After each transcription, the detector asks the classifier a choice question, "conversation" or "voicemail", about the whole transcript so far. Transcriptions that arrive while a question is in flight are asked about together in the next one.
  2. The detector keeps the latest answer but does not act on it while the caller is speaking.
  3. Once the caller has been quiet for decision_timeout and any question in flight has been answered, the latest answer becomes the verdict. The verdict is final.
  4. If the classifier never answered, because every call failed or timed out, the detector assumes a conversation, so a person is never left waiting on a bot that won’t speak.
On a conversation verdict, the gate releases the held speech and on_conversation_detected fires. On a voicemail verdict, the gate drops the held speech and the pipeline is interrupted. From then on the detector stops passing the other side’s frames downstream (only system, end, stop and agent lifecycle frames still flow), so the greeting never reaches the conversation LLM. After voicemail_response_delay of silence, on_voicemail_detected fires, once.

Event Handlers

Both handlers receive the detector, which is a frame processor, so a handler can push frames through it:

Metrics

The detector pushes its classifier’s on_metrics data into the pipeline as a MetricsFrame, so the time each classification took (and, with JevClassifier, the tokens it used) shows up alongside the pipeline’s other metrics.

Requirements

  • The detector is built for a cascaded pipeline: it reads the STT service’s transcriptions and holds the TTS service’s output.
  • An LLMClassifier needs a service that supports run_inference(), so a realtime (speech-to-speech) LLM cannot back it.

Removed in 1.12.0

The detector used to run a parallel pipeline with its own LLM. These parts of it were removed from pipecat.extensions.voicemail.voicemail_detector:
  • NotifierGate, ClassifierGate, ConversationGate and ClassificationProcessor
  • VoicemailDetector.CLASSIFIER_RESPONSE_INSTRUCTION and VoicemailDetector.DEFAULT_SYSTEM_PROMPT