Pipecat Voice Interface
Give our existing Letta agents natural, interruptible voice conversations across browser, phone, and local devices.
Superseded Added July 21, 2026. On September 20, EG chose the working dashboard voice stack as the production foundation, with Pipecat added incrementally behind its media interfaces. The current delivery order is in The Sunday Plan. The standalone sequence below remains as historical context.
Decision
Target architecture
Browser / phone / microphone
↓
Pipecat transport + VAD + speech-to-text
↓
LettaAgentProcessor (our adapter)
↓
Existing Letta agent + memory + tools
↓
Pipecat streaming text-to-speech
↓
Speaker / browser / phone
| Pipecat owns | Letta owns |
|---|---|
| WebRTC/WebSocket/phone transport | Agent identity and durable memory |
| Voice activity and turn detection | Reasoning, tools, and skills |
| STT, TTS, playback, and barge-in | Conversation history and state |
| Real-time media lifecycle and metrics | Long-running agent work |
First proof of concept
- Choose one existing Letta agent and preserve its current configuration unchanged.
- Build a small Python Pipecat service with a browser microphone client.
- Add a
LettaAgentProcessorthat sends a finalized transcript to the Letta messages API. - Translate only assistant-facing streamed text into Pipecat text frames; tool calls and internal steps must remain silent.
- Feed those frames into streaming TTS and support immediate playback interruption when the user begins speaking.
- Measure end-of-speech to first-audio latency and record transcripts for debugging.
Required adapter behavior
- Map each voice session to the correct Letta agent and conversation.
- Allow only one active request per agent conversation; queue or reject overlapping turns.
- Separate spoken assistant text from reasoning, tool calls, tool results, and status events.
- On barge-in, stop audio immediately and suppress every late token from the interrupted run.
- Cancel the Letta run when supported and safe; otherwise wait for it to terminate before submitting the next turn.
- Keep provider choices configurable so STT, TTS, and transports can be swapped without changing the adapter.
Implementation phases
Phase 1 — Local browser prototype
One agent, one user, browser audio, streaming STT/TTS, visible transcript, and basic interruption handling.
Phase 2 — Reliability
Run cancellation, stale-output protection, timeouts, reconnect behavior, structured logs, latency metrics, and automated adapter tests.
Phase 3 — Dashboard integration
Add a voice control to the dashboard Agents view, reuse the existing agent selector, and show listening/thinking/speaking/tool-working states.
Phase 4 — Additional transports
Evaluate phone, mobile, and scoreboard/DietPi clients only after the browser prototype is stable.
Risks and guardrails
- Do not duplicate Letta memory or agent tools inside Pipecat.
- Do not let intermediate tool output reach TTS.
- Require explicit confirmation before a voice command performs a destructive or externally visible action.
- Keep the initial prototype on a trusted local/Tailscale path; add authentication before broader exposure.
- Pin a Pipecat version during the prototype and isolate its adapter behind our own interface.
Prototype acceptance criteria
- The selected Letta agent remembers the same information in voice and text sessions.
- The user can interrupt speech and hear no stale continuation afterward.
- Tool calls execute through Letta, while only the final assistant response is spoken.
- Ten consecutive conversational turns complete without overlapping Letta requests.
- Disconnect and reconnect do not corrupt the agent conversation.
- The integration can change STT or TTS providers without changing Letta-specific code.
Before implementation
Re-check the current Pipecat and Letta streaming APIs, choose the first agent and transport, and write contract tests for event filtering, request serialization, interruption, and stale-token suppression before connecting live audio.