What is a UIWorker?
When you put a voice agent in front of an app, talking isn’t enough. The agent needs to see what the user sees and act on the screen: read the page, point at things, fill in fields, click buttons. AUIWorker is the server-side agent that makes this possible. It owns the screen, so the rest of your app doesn’t have to.
The connection is two-way, over the RTVI UI channel:
- Client → server. The client streams the screen to the
UIWorkeras accessibility snapshots, and forwards the user’s UI interactions as events. - Server → client. The
UIWorkerdrives the page back (scrolling, highlighting, selecting text, filling inputs, clicking, or running app-defined commands) and surfaces long-running work as progress cards.
UIWorker is the screen half of a voice/UI split:
- The voice agent holds the conversation and does all the talking. It never sees the page.
- The
UIWorkerowns the screen. When the voice agent needs something from it, it sends theUIWorkera job, and theUIWorkeranswers with short data: a label, a yes or no, a list of items. Never the whole page.
The two directions map to RTVI UI messages: the client sends
ui-snapshot and
ui-event; the UIWorker sends ui-command and ui-job-group. You rarely
touch these directly. PipelineWorker wires the channel up automatically when
RTVI is enabled (the default). See The RTVI
Standard for the wire protocol and the
UIWorker API reference for the
class.The two-way interface
What the UIWorker sees (client → server)
The screen, as a snapshot. The client sends an accessibility snapshot of the page whenever it changes, and theUIWorker keeps the latest one. Each element carries a stable ref the UIWorker uses to act on it. Rendered for an LLM, it looks like this:
snapshot property, and the text the user has selected, if any, through selection. When the UIWorker runs its own LLM, it injects the latest <ui_state> into that LLM’s context automatically.
User interactions, as events. The client dispatches app-defined events (a button click, a custom gesture) with sendUIEvent(name, payload). Route them to handlers with @ui_event(name); each runs in its own task:
What the UIWorker does (server → client)
Drives the page. TheUIWorker acts on the screen by sending UI commands. The built-in helpers cover the common actions, and send_command(name, payload) sends any app-defined command:
The standard client handlers ship in
@pipecat-ai/client-react; apps can override them or define their own command names.
Surfaces long work. Every job group the UIWorker dispatches shows up on the client as a progress card, with a line per peer agent. See Long-running work below.
Hello world
The smallestUIWorker is the class itself, with an LLM and a system prompt. It answers the built-in respond job: it runs one LLM turn with the latest <ui_state> in context, and the reply its LLM writes is the answer.
respond job and hands the answer back to its LLM, which speaks it:
UIWorker comes online to receive snapshots and jobs as soon as it starts:
1
Snapshot
The client streams the current screen as a
ui-snapshot. PipelineWorker
broadcasts it on the bus; the UIWorker keeps the latest one.2
Route
The user asks “what does the second story say?”. The voice LLM can’t see the
page, so it calls
ask_page, which sends a respond job to the UIWorker.3
Ground
The
UIWorker runs one LLM turn with the latest <ui_state> injected, so
its answer is grounded in what’s on screen.4
Speak
The job returns
{"answer": ...} to ask_page. The voice LLM phrases it
for the user, and its TTS speaks it.By default a
UIWorker is stateless: it clears its context at the start of
each respond job, so every turn sees only the current <ui_state> and
question. Set keep_history=True to accumulate history across turns, useful
for follow-ups like “and the one after that?”, at the cost of more tokens.Asking about the screen without an LLM turn
Arespond job runs a full LLM turn over the whole page. Many screen questions are smaller than that: which field is “the email field”? Is anything still unchecked? What has the user selected? The UIWorker answers these with its built-in screen job, using a classifier instead of an LLM turn.
screen_tools() gives the voice LLM a single screen(action, target, value) tool that sends this job:
Every answer is short data the voice LLM can use directly, such as
{"done": true, "label": "Email"}. Describe in the voice LLM’s prompt when to use each action. For example, a voice-guided form might say:
Choosing the classifier
Without aclassifier argument, the UIWorker answers screen questions with its own LLM through an LLMClassifier. That works, but costs an LLM call per question. Pass a JevClassifier for answers in about a tenth of a second, with calibrated probabilities:
Custom jobs
The built-in jobs are generic. When your app has its own vocabulary, such as a shopping list where the user adds, checks off and removes items, give yourUIWorker subclass its own @job handlers. A handler reads the screen with plain code, asks the classifier what it needs to, sends commands, and answers with a short result:
summary that lists what’s on the list from self.snapshot, lets the voice agent answer “what’s left?” from what’s actually on screen, including items the user checked off by hand.
The UIWorker also exposes the classifier questions the screen job uses, for your own handlers: which_element(), select_elements(), check_screen(), and act(), which finds an element and acts on it in one call:
Long-running work
When a job kicks off work that takes a while, fan it out to peer agents withjob_group(). Every group a UIWorker dispatches appears on the client as a cancellable progress card, with each agent’s progress streaming in as it arrives. Give the group a label to title the card:
Choosing an approach
They combine freely. One
UIWorker can answer respond and screen jobs and your own, and the voice LLM can hold screen_tools() alongside your app’s tools.
Migrating from earlier versions
Migrating from earlier versions
ReplyToolMixinis deprecated. Give the voice LLMscreen_tools()instead, or a custom job for app-specific actions, and let it say the answer.respond_to_job(text, tts_speak=True)is deprecated. AUIWorkershould not speak: answer with data, or let therespondjob answer with the LLM’s reply, and the voice LLM says it. -BaseUIWorkeris deprecated. Dispatch job groups from a@jobhandler on yourUIWorker, which reports them to the client itself.
What’s next
You’ve built agents that converse, call tools, and drive the screen. Next, learn how to transfer control between them.Agent Handoff
Activation, deactivation, and seamless control transfer
UIWorker API Reference
Full reference for
UIWorker, its built-in jobs, screen_tools, and UI
commands.