Overview
VoicemailDetector decides whether an outbound call reached a person or went to voicemail. It listens to what the other side says and asks a classifier. Until it has a verdict, its TTSGate holds the bot’s speech back, so a voicemail greeting is never talked over and a person hears the bot’s reply as soon as the verdict is in.
For a guide to setting it up, see Voicemail Detection.
Voicemail Detection Guide
Set up voicemail detection and leave a message
Example Implementation
A complete outbound bot with voicemail detection
Configuration
BaseClassifier
required
What decides between a person and a voicemail, such as a
JevClassifier or an
LLMClassifier. The detector owns
its lifecycle: it sets the classifier up and cleans it up with the pipeline.float
default:"1.0"
Seconds of silence after the caller stops speaking before the latest answer
decides. A greeting resumes after its pauses while a person stops and waits,
so a verdict on a fragment such as “hi, this is Sam” is not acted on until the
caller has really stopped.
float
default:"2.0"
Seconds of silence after a voicemail verdict before
on_voicemail_detected
fires, so the message is left after the greeting ends and the recording
starts. Speech, such as the rest of the greeting, restarts the wait.LLMService | None
default:"None"
deprecated
Deprecated since 1.12.0, removed in 2.0.0. An LLM service for the
classification. It is wrapped in an
LLMClassifier. Pass
classifier=LLMClassifier(llm=llm) instead.str | None
default:"None"
deprecated
Deprecated since 1.12.0, removed in 2.0.0. A system prompt for the
llm,
placed in front of the LLMClassifier’s own instructions. Pass an
LLMClassifier with your own instructions instead.Methods
detector
gate
TTSGate to place right after the TTS service. While the decision is pending it buffers TTS frames (TTSStartedFrame, TTSTextFrame, TTSAudioRawFrame, TTSStoppedFrame) and passes everything else. A conversation verdict releases the buffered frames in order. A voicemail verdict discards them, since they were meant for a person.
How the Decision Is Made
- After each transcription, the detector asks the classifier a choice question,
"conversation"or"voicemail", about the whole transcript so far. Transcriptions that arrive while a question is in flight are asked about together in the next one. - The detector keeps the latest answer but does not act on it while the caller is speaking.
- Once the caller has been quiet for
decision_timeoutand any question in flight has been answered, the latest answer becomes the verdict. The verdict is final. - If the classifier never answered, because every call failed or timed out, the detector assumes a conversation, so a person is never left waiting on a bot that won’t speak.
on_conversation_detected fires.
On a voicemail verdict, the gate drops the held speech and the pipeline is interrupted. From then on the detector stops passing the other side’s frames downstream (only system, end, stop and agent lifecycle frames still flow), so the greeting never reaches the conversation LLM. After voicemail_response_delay of silence, on_voicemail_detected fires, once.
Event Handlers
Both handlers receive the detector, which is a frame processor, so a handler can push frames through it:
Metrics
The detector pushes its classifier’son_metrics data into the pipeline as a MetricsFrame, so the time each classification took (and, with JevClassifier, the tokens it used) shows up alongside the pipeline’s other metrics.
Requirements
- The detector is built for a cascaded pipeline: it reads the STT service’s transcriptions and holds the TTS service’s output.
- An
LLMClassifierneeds a service that supportsrun_inference(), so a realtime (speech-to-speech) LLM cannot back it.
Removed in 1.12.0
The detector used to run a parallel pipeline with its own LLM. These parts of it were removed frompipecat.extensions.voicemail.voicemail_detector:
NotifierGate,ClassifierGate,ConversationGateandClassificationProcessorVoicemailDetector.CLASSIFIER_RESPONSE_INSTRUCTIONandVoicemailDetector.DEFAULT_SYSTEM_PROMPT