Transport plugin
Recommended.
pip install / npm i a plugin into your existing
livekit-agents worker. Your agent code, plugins, and provider keys do not
change. Only the transport changes. There is no LiveKit server at all.Front-proxy the media
The voice gateway terminates the caller. It bridges audio into your
existing LiveKit room through a sidecar. Use this path when you must
keep your LiveKit server in operation.
Browser compat shim
Preview. Change one import. A
livekit-client browser app then
publishes and subscribes over MoQT/QUIC instead of the SFU.These three paths solve different problems. The plugin replaces LiveKit’s
transport for a server-side agent. The front-proxy keeps your LiveKit
server and puts ClutchCall in front of it. The compat shim is for a
browser client. The paths combine. You can front-proxy a phone leg and
move a browser client onto the transport at the same time.
Path A — the transport plugin (recommended)
Install the plugin next to your agent. The plugin wrapsAgentSession.start.
It binds the session’s audio input and output to a QUIC/WebTransport connection
to ClutchCall instead of a LiveKit room. The LiveKit Agents SDK skips its own
RoomIO when a session already has custom audio IO set. Thus your entrypoint
runs verbatim: the same AgentSession, the same STT/LLM/TTS plugins, the
same provider keys, and the same tool calls. There is no room, no SFU, and no
WebRTC.
This is the same seam that LiveKit itself uses for its TCP console mode
(
TcpAudioInput / TcpAudioOutput). The plugin points that seam at our QUIC
edge.:443 and opens no
inbound ports. Thus it works behind NAT and needs no public address. One
session carries three lanes:
- presence/registry (liveness)
- audio over QUIC datagrams
- per-call control (job assign, flush, barge-in, playback finished) over a WebTransport stream
Install
- Python
- TypeScript
livekit-agents>=1.0. The
quic extra installs aioquic (the WebTransport client). Without it, the
package still imports, but it cannot connect.Run your agent over the transport
Replace your worker’s entrypoint runner (cli.run_app(...) /
cli.runApp(...)) with ours. The code it calls, your entrypoint, does not
change.
- Python
- TypeScript
worker.py
run_agent_worker does these tasks:
- It authenticates one connection.
- It registers the worker under
agent_name. - It serves the calls that the engine assigns to it. For each call, it starts
your
entrypointwith aJobContextbound to that call’s media.
Point a ClutchCall agent at the worker
The engine side models a remote agent as aVENDOR_BRIDGE node with
vendor_provider set to clutchcall. This means external agent
orchestration, and ClutchCall stays the transport. vendor_room is the
agent_name that your worker registered.
agent-config.json
With this config, the runtime runs no turn detection and no model in
the engine. It sends the caller’s audio to your worker. It streams your agent’s
audio back onto the call. Attach the agent to a number or trunk in the same way
as a native agent. See
SIP Trunking and
Number Provisioning.
What your agent sees
Yourentrypoint receives a JobContext. Its ctx.room is a veneer, not
a LiveKit room. The veneer carries the fields that agents read:
Caller and called numbers surface where LiveKit’s SIP integration puts them —
but the two plugins put them in different places today:
- Python
- TypeScript
agents.updateConfig with { patch: { metadata: { env: "prod", team: "sales" } } }). The engine forwards them on every call dispatched to your
worker, and the plugin surfaces them where LiveKit puts dispatch metadata:
Tool calls and token analytics
Your agent runs its own tool loop and its own STT/LLM/TTS on your provider keys, so the engine can’t see either — but you can forward both over the control stream and they land on the ClutchCall timeline (and fire thevoice.agent.tool_called / voice.agent.usage webhooks) alongside native
agents. Wire them from your agent’s hooks:
sendUsage, sendToolCall).
Redact anything sensitive yourself — args, results and usage fields are
forwarded as-is (the engine can’t see inside your agent to redact for you).
Native ClutchCall agents get tool logging automatically; usage metering for
native agents is on the roadmap.
Outbound calling (single + batch)
The same worker serves calls the platform dials out. Trigger dials with the LiveKit-shaped API in the plugin — it wrapstelephony.originate on the
administration API, authenticated with an org-scoped mpk_… key (mint one
under Settings → API keys with scope mcp:telephony:write; set it as
CLUTCHCALL_API_KEY alongside your org id):
new ClutchCallAPI(),
api.sip.createSipParticipant(...), api.sip.createSipCampaign(...)). A
runnable example lives in the starter repo as src/outbound.py.
Campaign pacing, limits, control (list_campaigns / abort_campaign), and
what the worker receives on answer are covered in full in
Outbound Calls from a LiveKit Agent.
For the rest of the administration surface use the
Admin API SDK.
Reaching the agent from a browser
The same worker also serves web callers.@clutchcall/agent-widget is an
embeddable widget whose media runs on MoQT/QUIC. The engine republishes the
browser’s mic track onto your worker’s media plane and returns the reply on the
same downlink, so your agent code sees a web caller exactly as it sees a phone
caller. See
Embed an Agent Widget on Your Website.
Camera and keyboard input
A browser caller can also send their camera and their keyboard, and your worker receives either one — but only if it asks. Declare what you accept and read it:- Python
- TypeScript
worker.py
my_agent.py
Why a callback and not
session.input.video. LiveKit’s session wants
decoded rtc.VideoFrames for a continuous track. These are stills, so feeding
that seam would mean shipping a decoder to undo work nobody did. Decode them
yourself if you want them on the session.- Default off. A worker that does not ask gets no frames, and the engine does not subscribe the track at all — an audio-only agent costs exactly what it did before.
- The caller is told the truth. The engine announces
video/textfor the call from what the worker that took it declared, so the widget shows a control only when pressing it does something. PSTN callers have neither a camera nor a keyboard, so both are browser-caller features.
Recordings, CDR, and transcripts
Recording is a call-layer capture, so a call your worker handles is recorded exactly like a native agent’s — both legs, same storage, same API. Fetch the audio with the Recordings API, keyed by thecall_sid your agent sees as ctx.job.id.
Current limits
Python and TypeScript only
Python and TypeScript only
livekit-plugins-clutchcall (Python) and @clutchcall/livekit-transport (TypeScript) track the
two LiveKit Agents SDKs. There is no Go or Rust plugin.Audio is pcm16 today
Audio is pcm16 today
The wire codec is raw little-endian
pcm16. Opus fits into the same codec
seam without changes to the session or IO classes, but it is not connected
yet. Thus budget bandwidth for uncompressed audio between your worker and
the edge.Register-and-wait is the shipped launch mode
Register-and-wait is the shipped launch mode
The worker holds a connection open, and the engine dispatches calls to it.
This mode is correct for always-on workers on a VM or in Kubernetes, and it
is warm on the first call. A per-call HTTP trigger is not yet available.
That trigger would serve scale-to-zero serverless workers that dial back on
demand.
Audio needs QUIC egress
Audio needs QUIC egress
The worker needs outbound UDP/443 to the edge. The WebSocket fallback
covers presence only. Media over TCP reintroduces head-of-line blocking and
is a degraded last resort. Agents run on server infrastructure, where QUIC
egress is the norm.
Path B — front-proxy the media
Use this path when you must keep your LiveKit server in operation. For example, other participants join the room, or another part of your stack depends on it. In this path, the gateway owns the caller edge: a PSTN SIP trunk, or a browser leg over QUIC on the single:443 plane. The gateway
decodes the edge to the runtime’s 8 kHz PCM audio bus. It relays that audio
over a WebSocket media bridge to a bridge sidecar that you run next to your
LiveKit deployment. That sidecar joins your room as an ordinary participant.
What is shipped and what you deploy. The gateway-side bridge is in-tree
and shipped. It consists of a vendor-bridge pipeline node plus the WebSocket
media handler. The bridge sidecar that joins your LiveKit room is a
component that you run next to your LiveKit deployment. It uses your LiveKit
URL, API key, and secret.
VENDOR_BRIDGE shape. Here vendor_provider
names the vendor instead of ClutchCall:
agent-config.json
Path C — the browser compat shim (preview)
If your LiveKit surface is a browser app (livekit-client), you can run
the client’s audio over our transport directly. The livekit-compat shim
re-implements the livekit-client API (Room, RoomEvent, tracks,
participants) on top of MoQT/QUIC. The common voice path is a one-line import
change. That path covers these actions: publish the mic, subscribe to remote
audio, and exchange data. There is no SFU, and the shim never uses WebRTC’s
transport.
RTCRtpScriptTransform). It refuses to
connect on a browser that would silently fall back to plaintext WebRTC. Today
that means Chromium-family browsers. The full deep-dive is in
Keep Your Existing Runtime and
WebRTC Diversion.
Which path?
Related
Run a LiveKit Agent on the Transport
The end-to-end recipe: install the plugin, register a worker, and take a call.
Keep Your Existing Runtime
The full transport-only spectrum: the plugin, the shim, and raw MoQT.
From Self-Hosted LiveKit
The end-to-end migration path: trunks, tokens, and rollout.
Custom Agent Runtime
The same patterns for any non-LiveKit runtime.
Outbound Calls from a LiveKit Agent
Dial one number or run a paced campaign. The same worker answers.
Embed an Agent Widget
Put the same agent on your website. One tag plus one backend route.
Recordings API
List recordings, presign a download, or get pushed a webhook.

