Voice Input for the Agent Dashboard

A TDD + GoF-pattern plan to add browser push-to-talk voice capture to each agent's Input Options panel — reusing lettabot's already-working local whisper.cpp pipeline.

Source spec: dashboard/audio_input/audio_input.md  ·  Target page: Agent detail → Input Options tab  ·  Date: 2026-06-06

Strategy Adapter Factory Pipeline State (front-end)

Build status — 2026-06-06: shipped, working end-to-end on Android

The full loop — phone mic → HTTPS upload → local whisper.cpp → cleanup agent → deliver to the chosen agent — works from the phone today. One known accuracy issue remains (associate-name mixup, see below). Committed to git so the team is in sync.

One-line summary: Press Start → record in the browser → on Stop, upload the audio to the dashboard server → transcribe locally with whisper.cpp → clean the transcript with a small, fast Letta agent (so Friday becomes Frita) → drop the cleaned text into the message box and send it to the chosen agent.

1. What already exists (so we don't reinvent it)

PieceWhere it livesStatus
Browser record / stop UI (MediaRecorder, webm upload)planner/new_voice_menu_app/templates/index.htmlWorking — we copy the interaction model
Local whisper.cpp transcription recipelettabot/src/transcription/whispercpp.tsWorking in lettabot today — we port the recipe to Python
whisper-cli binary~/whisper.cpp/build/bin/whisper-cliPresent
Model~/whisper.cpp/models/ggml-base.en.bin (148 MB)Present
ffmpeg (audio → 16k mono wav)lettabot/.venv_ffmpeg/.../imageio_ffmpeg/binaries/ffmpeg-linux-x86_64-v7.0.2Present (no system ffmpeg)
Dashboard server + agent send (/api/test)dashboard/server.pyWorking — we reuse it for delivery
Known agent names (for name correction)LETTA_AGENTS in server.pyScissari, Frita, Hailey, Jeri, Mazda

2. The pipeline

browser dashboard server (server.py) Letta API ─────── ──────────────────────────── ───────── [Start] ── getUserMedia ──┐ │ MediaRecorder │ [Stop] ── webm Blob ───────┼──▶ POST /api/voice (raw body, X-Agent-Id) │ │ │ ▼ VoicePipeline.process(bytes) │ ┌─ TranscriptionStrategy (WhisperCpp) ─┐ │ │ ffmpeg → 16k mono wav │ │ │ whisper-cli -otxt -nt → transcript │ raw: "Tell Friday about this." │ └──────────────────────────────────────┘ │ │ │ ▼ CleanupStrategy (LettaAgentCleanup) ──▶ POST /v1/agents//messages │ cleaned: "Tell Frita about this." ◀──────── (gemini-2.5-flash-lite) │ │ ◀────────┴─ JSON { raw_transcript, cleaned_text } fill #am-test-text with cleaned_text auto-trigger existing [Send] ─────────▶ POST /api/test ─────────────────▶ main agent

The cleanup step is deliberately a separate, cheaper model in front of the main agent. whisper hears Friday; the cleanup agent — primed with the list of real agent names — rewrites it to Frita before the main agent ever sees it.

3. Design — GoF patterns, used where they earn their place

PatternApplied toWhy
StrategyTranscriptionStrategy → WhisperCppTranscriber; CleanupStrategy → LettaAgentCleanupTranscription engine and cleanup brain are both swappable (whisper.cpp today, OpenAI later; Letta agent today, anything later) without touching the pipeline.
AdapterLettaClient wrapping the HTTP callsCleanup and main-send share one thin client; tests inject a fake instead of hitting the network.
Factorybuild_transcriber() / build_cleanup() reading configserver.py stays declarative; tests substitute fakes by config.
PipelineVoicePipeline.process(bytes) → {raw, cleaned}Composes transcribe → cleanup; falls back to the raw transcript if cleanup fails so a cleanup hiccup never blocks you.
State (front-end)Tiny idle ↔ recording ↔ processing machine for the buttonOne source of truth for button colour/label, the LED indicator, and the cycling "Recording…" text.

4. File layout

dashboard/
  voice/
    __init__.py
    config.py          # paths + cleanup-agent id from env, with the lettabot defaults baked in
    letta_client.py    # Adapter — GET/POST helpers around the Letta API
    transcription.py   # TranscriptionStrategy, WhisperCppTranscriber, build_transcriber()
    cleanup.py         # CleanupStrategy, LettaAgentCleanup, build_cleanup()
    pipeline.py        # VoicePipeline.process(audio_bytes, filename) -> {raw, cleaned}
  tests/
    conftest.py
    test_transcription.py   # ffmpeg/whisper args + .txt read; missing-binary/model errors (subprocess mocked)
    test_cleanup.py         # posts to cleanup agent; Friday→Frita via fake client; raw-fallback on failure
    test_pipeline.py        # composition with injected fakes; cleanup-failure → raw fallback
    test_endpoints.py       # POST /api/voice returns correct JSON with a fake pipeline
  requirements-dev.txt      # pytest
  server.py            # + POST /api/voice ; + ThreadingHTTPServer (see Q2)
  dashboard.html       # rename tab → "Input Options"; add Start/Stop button + recording indicator

5. Whisper recipe (ported verbatim from lettabot)

# 1. browser blob (audio/webm;codecs=opus) written to a temp file
ffmpeg -y -i source.webm -ar 16000 -ac 1 -c:a pcm_s16le input.wav
# 2. transcribe
whisper-cli -m ggml-base.en.bin -f input.wav -l auto -of transcript -otxt -nt
# 3. read transcript.txt → strip → done

Defaults resolved from env with lettabot's exact paths as fallback: WHISPER_CPP_BIN, WHISPER_MODEL_PATH, FFMPEG_BIN.

6. UI changes (dashboard.html)

7. TDD order (pytest — failing tests first)

  1. test_transcription — builds the correct ffmpeg + whisper-cli args, reads the .txt; error paths for missing binary/model (subprocess mocked, no real audio needed).
  2. test_cleanup — posts to the cleanup agent, parses the reply; Friday → Frita via a fake Letta client; raw-fallback when cleanup errors.
  3. test_pipeline — composes transcribe → cleanup with injected fakes; cleanup failure yields the raw transcript.
  4. test_endpoints — POST /api/voice returns the right JSON shape with a fake pipeline (no whisper, no network).

8. Open questions for the team

Decisions locked — 2026-06-06

#Decision
Q1Tailscale Serve — front the dashboard with a real HTTPS cert on the MagicDNS name (https://desktop-2obsqmc-24.tailb8fc54.ts.net/) via tailscale serve --bg 8765. The phone (samsung-sm-s156v) is already on the tailnet. Secure context ⇒ mic works, no cert warnings. Self-signed HTTPS is the fallback only if the tailnet's "HTTPS Certificates" feature can't be enabled. Raw http:// over the tailnet IP does not work — still an insecure context.
Q2Switch server.py to ThreadingHTTPServer so transcription doesn't freeze the dashboard.
Q3Clear the cleanup agent's own message history on each call (Letta messages/clear, same as /api/test). Only its scratch conversation is cleared — nothing user-facing.
Q4Add a toggle: Auto Send ↔ Review then Send. Auto Send delivers the cleaned text to the agent immediately; Review then Send drops it in the box for a manual Send.
Q5Create transcript-cleanup-agent on gemini-2.5-flash-lite (confirmed).
Q6Stay on ggml-base.en (English) + cleanup for name fixes.
Q7Body text → "Meeting with {name}" (replaces the current "Chat with {name}:").

Q1 — Microphone over the network on Android (highest risk)

The dashboard is served over plain HTTP on 0.0.0.0:8765. A phone reaches it at http://<machine-ip>:8765. Browsers — including Android Chrome — block getUserMedia() (the mic) outside a secure context (https:// or localhost). So voice will work on the host's own localhost but silently fail on the phone until we pick one of:

  1. HTTPS with a self-signed cert on the dashboard (accept the warning once per device). Self-contained, no extra services.
  2. A tunnel (cloudflared / ngrok) that gives a real https:// URL. Easiest trust story, adds a dependency.
  3. Per-device override: add the origin to chrome://flags/#unsafely-treat-insecure-origin-as-secure on each phone. Zero code, but manual per device.

Recommendation: option 1 (self-signed HTTPS) for the dashboard. Which do you want?

Q2 — Make the server threaded

server.py runs a single-threaded HTTPServer. A whisper transcription takes a few seconds and would freeze the entire dashboard (the 3-second pollers stall) while it runs. Proposed fix: switch to ThreadingHTTPServer (a one-line change). Any objection to that touching shared server.py?

Q3 — Keep the cleanup agent stateless

Letta agents accumulate history. If we reuse one cleanup agent, every transcript piles into its context (noisy, slowly more expensive). The existing /api/test already calls messages/clear before each send. Plan: clear the cleanup agent's history on each call so it behaves statelessly. OK?

Q4 — Auto-send, or let me review first?

The spec says Stop → transcribe → cleanup → "pass it to the main Agent." Two flavours:

Recommendation: fill the box and auto-send, but show the cleaned text first so nothing is hidden. Good, or do you want a manual confirm?

Q5 — Cleanup agent + model

Plan: create a dedicated transcript-cleanup-agent on gemini-2.5-flash-lite (cheapest/fastest live model), system prompt = "fix speech-to-text errors and agent names only, don't change meaning," with the known agent list baked in. Confirm the name & model, or point me at a different one.

Q6 — English-only model vs multilingual

ggml-base.en.bin is English-only and fast. Agent names like Frita, Scissari may be mis-heard — but that's exactly what the cleanup step fixes. Stay on base.en + cleanup, or download a multilingual model? Recommendation: stay on base.en.

Q7 — Page body wording (minor)

The spec quotes the panel as saying Send a test message to {Agent_Name}, but the live page actually reads Chat with {name}:. Per "keep the existing text the same" I'll leave it as-is and only rename the tab. Say the word if you'd rather I switch the body text to the spec's wording.

9. Status & what's next

This started as a plan; it has shipped. The pytest suite is green (21), the UI is wired, the cleanup agent exists, and the full loop works end-to-end from an Android phone over the Tailscale HTTPS URL (Q1 resolved — secure context obtained). All Q-blocks below were the original open questions and are now answered; the Decisions locked table in §8 records the choices.

The last accuracy item — the associate/agent name mixup — has a fix in place at both the transcription stage (whisper prompt-priming) and the cleanup stage (the cleanup agent's name-correction), and a teammate's Claude agent may have already resolved it; treat it as likely fixed, pending a final spoken-name spot check. voice_transcripts.json records raw-vs-cleaned output for any lingering mishears; the remaining lever, if needed, is a deterministic fuzzy name-match as a final pass.

— Q-blocks below are kept for history; see §8 for the locked decisions.