Skip to main content

Overview

When a bot places an outbound call, it needs to know who picked up. A person expects a quick, natural reply. A voicemail greeting should be listened to in full, and then the bot should leave a message after the beep. VoicemailDetector makes that call. It listens to what the other side says and asks a classifier whether it’s a person or a voicemail. Meanwhile the bot’s reply is generated as usual but held back, so a person hears it as soon as the verdict is in, and a voicemail greeting is never talked over.

How It Works

The detector is two processors: the detector itself, after the STT service, and a gate, right after the TTS service.
  1. Listen. After each transcription, the detector asks its classifier whether the transcript so far sounds like a person or a voicemail. The classifier runs in the background, so frames keep flowing.
  2. Hold. While the decision is pending, the gate buffers the bot’s speech. The LLM and TTS run as normal, so the reply is ready the moment the verdict arrives.
  3. Wait for silence. No answer is acted on while the caller is speaking. “Hi, this is Sam” is what a person says when they pick up, and also how a greeting starts. Only what follows tells them apart: a greeting keeps going after its pauses, while a person stops and waits. Once the caller has been quiet for decision_timeout, the latest answer becomes the verdict.
  4. Act. For a conversation, the gate releases the held speech and the call goes on. For a voicemail, the held speech is dropped, and once the greeting has been quiet for voicemail_response_delay, on_voicemail_detected fires so you can leave a message.

Basic Setup

1. Create the Detector

Pass the detector a classifier. JevClassifier answers in about a tenth of a second, which matters here: every bit of classification time is time a person waits in silence after saying “hello?”.
JevClassifier needs the jev extra (uv add "pipecat-ai[jev]") and a TypeSafe API key. To classify with an LLM you already use instead, see Using an LLM Classifier.

2. Configure the Pipeline

The detector needs two components in your pipeline:
  • detector(): between the STT service and the user context aggregator
  • gate(): immediately after the TTS service

3. Leave a Message

When the call reaches voicemail, on_voicemail_detected fires once the greeting has finished. The handler receives the detector, which is a frame processor, so you can push frames through it. For example, speak a message and then end the call:
After a voicemail verdict, the detector stops passing the other side’s speech to the rest of the pipeline, so your conversation LLM never replies to the greeting. The handler fires once: a beep or a prompt heard while the message plays doesn’t trigger it again.

Detecting a Conversation

When a person answers, there’s nothing to do: the gate releases the bot’s held reply and the conversation continues. If you want to react, for example to log the outcome or start a timer, handle on_conversation_detected:
If the classifier can’t answer at all, because it failed or timed out, the detector assumes a conversation once the caller goes quiet. A person is never left waiting for a bot that won’t speak.

Configuration Options

Response Timing

Two parameters control the timing. Tune them against the greetings your calls actually reach. decision_timeout (default 1 second) is how long the caller must be quiet before the latest answer decides. Raise it if greetings with long pauses get classified as conversations. Lower it to answer people faster.
voicemail_response_delay (default 2 seconds) is how long the greeting must be quiet after a voicemail verdict before on_voicemail_detected fires. It ensures the greeting has finished and the recording has started before you speak.

Using an LLM Classifier

To classify with an LLM service instead of Jev, wrap it in an LLMClassifier. Any service that supports run_inference() works, and a small, fast model is usually enough. It doesn’t need to be the LLM in your pipeline:
An LLM request takes longer than a Jev call, and that time adds to the silence a person hears before the bot replies.

Migrating from the llm Parameter

Before 1.12.0, the detector took an llm and an optional custom_system_prompt. Both are deprecated and will be removed in 2.0.0. Until then, an llm is wrapped in an LLMClassifier for you.
A custom_system_prompt no longer needs to ask for a CONVERSATION or VOICEMAIL reply; that constant, CLASSIFIER_RESPONSE_INSTRUCTION, has been removed. If you customized the prompt, move your guidance into an LLMClassifier’s instructions, and keep its default instructions for the reply format:
The event handlers now receive the detector itself rather than an internal processor. Code that pushes frames from a handler keeps working.

Next Steps

Try the Voicemail Detection Example

A complete bot that detects voicemail, leaves a message, and ends the call.

VoicemailDetector Reference

Parameters, events, and how the decision is made

Pipecat Classifiers

How classifiers work and how to choose one