Letta Agents — Original Proposal

Pipecat Voice Interface

Give our existing Letta agents natural, interruptible voice conversations across browser, phone, and local devices.

Superseded Added July 21, 2026. On September 20, EG chose the working dashboard voice stack as the production foundation, with Pipecat added incrementally behind its media interfaces. The current delivery order is in The Sunday Plan. The standalone sequence below remains as historical context.

Decision

Current direction: extend the dashboard stack through interfaces. Letta remains the agent brain and source of truth for memory, identity, tools, skills, and conversation state. Pipecat can supply real-time media features through adapters as each feature is integrated and verified.

Target architecture

Browser / phone / microphone
          ↓
Pipecat transport + VAD + speech-to-text
          ↓
LettaAgentProcessor (our adapter)
          ↓
Existing Letta agent + memory + tools
          ↓
Pipecat streaming text-to-speech
          ↓
Speaker / browser / phone
Pipecat ownsLetta owns
WebRTC/WebSocket/phone transportAgent identity and durable memory
Voice activity and turn detectionReasoning, tools, and skills
STT, TTS, playback, and barge-inConversation history and state
Real-time media lifecycle and metricsLong-running agent work

First proof of concept

  1. Choose one existing Letta agent and preserve its current configuration unchanged.
  2. Build a small Python Pipecat service with a browser microphone client.
  3. Add a LettaAgentProcessor that sends a finalized transcript to the Letta messages API.
  4. Translate only assistant-facing streamed text into Pipecat text frames; tool calls and internal steps must remain silent.
  5. Feed those frames into streaming TTS and support immediate playback interruption when the user begins speaking.
  6. Measure end-of-speech to first-audio latency and record transcripts for debugging.

Required adapter behavior

Implementation phases

Phase 1 — Local browser prototype

One agent, one user, browser audio, streaming STT/TTS, visible transcript, and basic interruption handling.

Phase 2 — Reliability

Run cancellation, stale-output protection, timeouts, reconnect behavior, structured logs, latency metrics, and automated adapter tests.

Phase 3 — Dashboard integration

Add a voice control to the dashboard Agents view, reuse the existing agent selector, and show listening/thinking/speaking/tool-working states.

Phase 4 — Additional transports

Evaluate phone, mobile, and scoreboard/DietPi clients only after the browser prototype is stable.

Risks and guardrails

Interruption is the critical seam. Stopping TTS does not automatically stop a Letta run. Every run needs an ID or generation token so late output can never leak into the next spoken turn.

Prototype acceptance criteria

Before implementation

Re-check the current Pipecat and Letta streaming APIs, choose the first agent and transport, and write contract tests for event filtering, request serialization, interruption, and stale-token suppression before connecting live audio.