A TDD + GoF-pattern plan to add browser push-to-talk voice capture to each agent's Input Options panel — reusing lettabot's already-working local whisper.cpp pipeline.
Strategy Adapter Factory Pipeline State (front-end)
Build status — 2026-06-06: shipped, working end-to-end on Android
dashboard/voice/ package (Strategy / Adapter / Factory / Pipeline) — 21 pytest tests green (so venv → python -m pytest dashboard/tests/).transcript-cleanup-agent live on gemini-2.5-flash-lite; whisper is also primed with the agent-name list so mishears are reduced before cleanup.server.py: ThreadingHTTPServer + POST /api/voice live on :8765; raw+cleaned transcripts logged to voice_transcripts.json for diagnosing mishears.samsung-sm-s156v.--prompt "Agent names: Scissari, Frita, Hailey, Jeri, Mazda." so a name is heard correctly in the first place (e.g. "Mazda" no longer becomes "Melissa"). See voice/config.py & voice/transcription.py.transcript-cleanup-agent is instructed to "correct agent names to the closest known agent," with the same name list (voice/cleanup.py).voice_transcripts.json logs raw-vs-cleaned for diagnosis; next lever would be a deterministic fuzzy name-match as a final pass.One-line summary: Press Start → record in the browser → on Stop, upload the audio to the dashboard server → transcribe locally with whisper.cpp → clean the transcript with a small, fast Letta agent (so Friday becomes Frita) → drop the cleaned text into the message box and send it to the chosen agent.
| Piece | Where it lives | Status |
|---|---|---|
Browser record / stop UI (MediaRecorder, webm upload) | planner/new_voice_menu_app/templates/index.html | Working — we copy the interaction model |
| Local whisper.cpp transcription recipe | lettabot/src/transcription/whispercpp.ts | Working in lettabot today — we port the recipe to Python |
whisper-cli binary | ~/whisper.cpp/build/bin/whisper-cli | Present |
| Model | ~/whisper.cpp/models/ggml-base.en.bin (148 MB) | Present |
| ffmpeg (audio → 16k mono wav) | lettabot/.venv_ffmpeg/.../imageio_ffmpeg/binaries/ffmpeg-linux-x86_64-v7.0.2 | Present (no system ffmpeg) |
Dashboard server + agent send (/api/test) | dashboard/server.py | Working — we reuse it for delivery |
| Known agent names (for name correction) | LETTA_AGENTS in server.py | Scissari, Frita, Hailey, Jeri, Mazda |
The cleanup step is deliberately a separate, cheaper model in front of the main agent. whisper hears Friday; the cleanup agent — primed with the list of real agent names — rewrites it to Frita before the main agent ever sees it.
| Pattern | Applied to | Why |
|---|---|---|
| Strategy | TranscriptionStrategy → WhisperCppTranscriber; CleanupStrategy → LettaAgentCleanup | Transcription engine and cleanup brain are both swappable (whisper.cpp today, OpenAI later; Letta agent today, anything later) without touching the pipeline. |
| Adapter | LettaClient wrapping the HTTP calls | Cleanup and main-send share one thin client; tests inject a fake instead of hitting the network. |
| Factory | build_transcriber() / build_cleanup() reading config | server.py stays declarative; tests substitute fakes by config. |
| Pipeline | VoicePipeline.process(bytes) → {raw, cleaned} | Composes transcribe → cleanup; falls back to the raw transcript if cleanup fails so a cleanup hiccup never blocks you. |
| State (front-end) | Tiny idle ↔ recording ↔ processing machine for the button | One source of truth for button colour/label, the LED indicator, and the cycling "Recording…" text. |
dashboard/
voice/
__init__.py
config.py # paths + cleanup-agent id from env, with the lettabot defaults baked in
letta_client.py # Adapter — GET/POST helpers around the Letta API
transcription.py # TranscriptionStrategy, WhisperCppTranscriber, build_transcriber()
cleanup.py # CleanupStrategy, LettaAgentCleanup, build_cleanup()
pipeline.py # VoicePipeline.process(audio_bytes, filename) -> {raw, cleaned}
tests/
conftest.py
test_transcription.py # ffmpeg/whisper args + .txt read; missing-binary/model errors (subprocess mocked)
test_cleanup.py # posts to cleanup agent; Friday→Frita via fake client; raw-fallback on failure
test_pipeline.py # composition with injected fakes; cleanup-failure → raw fallback
test_endpoints.py # POST /api/voice returns correct JSON with a fake pipeline
requirements-dev.txt # pytest
server.py # + POST /api/voice ; + ThreadingHTTPServer (see Q2)
dashboard.html # rename tab → "Input Options"; add Start/Stop button + recording indicator
# 1. browser blob (audio/webm;codecs=opus) written to a temp file
ffmpeg -y -i source.webm -ar 16000 -ac 1 -c:a pcm_s16le input.wav
# 2. transcribe
whisper-cli -m ggml-base.en.bin -f input.wav -l auto -of transcript -otxt -nt
# 3. read transcript.txt → strip → done
dashboard.html)Chat Interface → Input Options. Keep the existing chat text & layout per the spec..am-btn: default green background, black "Start" text.Recording / Recording. / Recording.. / Recording....txt; error paths for missing binary/model (subprocess mocked, no real audio needed).Friday → Frita via a fake Letta client; raw-fallback when cleanup errors.POST /api/voice returns the right JSON shape with a fake pipeline (no whisper, no network).Decisions locked — 2026-06-06
| # | Decision |
|---|---|
| Q1 | Tailscale Serve — front the dashboard with a real HTTPS cert on the MagicDNS name (https://desktop-2obsqmc-24.tailb8fc54.ts.net/) via tailscale serve --bg 8765. The phone (samsung-sm-s156v) is already on the tailnet. Secure context ⇒ mic works, no cert warnings. Self-signed HTTPS is the fallback only if the tailnet's "HTTPS Certificates" feature can't be enabled. Raw http:// over the tailnet IP does not work — still an insecure context. |
| Q2 | Switch server.py to ThreadingHTTPServer so transcription doesn't freeze the dashboard. |
| Q3 | Clear the cleanup agent's own message history on each call (Letta messages/clear, same as /api/test). Only its scratch conversation is cleared — nothing user-facing. |
| Q4 | Add a toggle: Auto Send ↔ Review then Send. Auto Send delivers the cleaned text to the agent immediately; Review then Send drops it in the box for a manual Send. |
| Q5 | Create transcript-cleanup-agent on gemini-2.5-flash-lite (confirmed). |
| Q6 | Stay on ggml-base.en (English) + cleanup for name fixes. |
| Q7 | Body text → "Meeting with {name}" (replaces the current "Chat with {name}:"). |
The dashboard is served over plain HTTP on 0.0.0.0:8765. A phone reaches it at http://<machine-ip>:8765. Browsers — including Android Chrome — block getUserMedia() (the mic) outside a secure context (https:// or localhost). So voice will work on the host's own localhost but silently fail on the phone until we pick one of:
https:// URL. Easiest trust story, adds a dependency.chrome://flags/#unsafely-treat-insecure-origin-as-secure on each phone. Zero code, but manual per device.Recommendation: option 1 (self-signed HTTPS) for the dashboard. Which do you want?
server.py runs a single-threaded HTTPServer. A whisper transcription takes a few seconds and would freeze the entire dashboard (the 3-second pollers stall) while it runs. Proposed fix: switch to ThreadingHTTPServer (a one-line change). Any objection to that touching shared server.py?
Letta agents accumulate history. If we reuse one cleanup agent, every transcript piles into its context (noisy, slowly more expensive). The existing /api/test already calls messages/clear before each send. Plan: clear the cleanup agent's history on each call so it behaves statelessly. OK?
The spec says Stop → transcribe → cleanup → "pass it to the main Agent." Two flavours:
Recommendation: fill the box and auto-send, but show the cleaned text first so nothing is hidden. Good, or do you want a manual confirm?
Plan: create a dedicated transcript-cleanup-agent on gemini-2.5-flash-lite (cheapest/fastest live model), system prompt = "fix speech-to-text errors and agent names only, don't change meaning," with the known agent list baked in. Confirm the name & model, or point me at a different one.
ggml-base.en.bin is English-only and fast. Agent names like Frita, Scissari may be mis-heard — but that's exactly what the cleanup step fixes. Stay on base.en + cleanup, or download a multilingual model? Recommendation: stay on base.en.
The spec quotes the panel as saying Send a test message to {Agent_Name}, but the live page actually reads Chat with {name}:. Per "keep the existing text the same" I'll leave it as-is and only rename the tab. Say the word if you'd rather I switch the body text to the spec's wording.
This started as a plan; it has shipped. The pytest suite is green (21), the UI is wired, the cleanup agent exists, and the full loop works end-to-end from an Android phone over the Tailscale HTTPS URL (Q1 resolved — secure context obtained). All Q-blocks below were the original open questions and are now answered; the Decisions locked table in §8 records the choices.
The last accuracy item — the associate/agent name mixup — has a fix in place at both the transcription stage (whisper prompt-priming) and the cleanup stage (the cleanup agent's name-correction), and a teammate's Claude agent may have already resolved it; treat it as likely fixed, pending a final spoken-name spot check. voice_transcripts.json records raw-vs-cleaned output for any lingering mishears; the remaining lever, if needed, is a deterministic fuzzy name-match as a final pass.