SwarmForge's six-role pipeline did land a working fix (94% mutation coverage, 20/20 acceptance scenarios, all six worktrees converged on QA's final commit). But it took direct human/supervisor intervention at least 12 distinct times to get there: a silently-dead daemon, a role that skipped notifying two of its teammates, three separate instances of a role ignoring or diluting an explicit instruction, a stale status display masking real progress, an OOM cascade that killed unrelated production services, and a handoff format constraint nobody had hit before. None of these were fatal individually — the run still finished — but none of them were caught by SwarmForge itself either. Every catch below came from a human (or a human-directed supervisor) checking the actual filesystem/process state instead of trusting a role's self-report.
Six roles, each its own git worktree and its own long-running claude CLI session, coordinate by dropping files into each other's .swarmforge/handoffs/inbox — moved by a background daemon, handoffd. There is no central scheduler; every role decides for itself when to check its inbox and what to do next. The diagram below marks every failure point found this run.
handoffd is wrapped in a sleep-inhibitor (systemd-inhibit --what=sleep:idle) so the box doesn't nap mid-run. Under WSL, systemd-inhibit fails non-interactively even when systemctl is-system-running reports running — polkit wants an interactive prompt that never comes. The daemon's own startup gate passes, then the actual inhibitor call fails, and the daemon never starts. Two roles (coder→cleaner, hardender→QA) had correctly written handoffs to disk — nothing was lost — but nothing was moving them.
swarmforge.bb's sleep-inhibitor-prefix gates on system-running state, not on whether the inhibitor call itself succeeds. A failed inhibitor call silently prevented the wrapped daemon from ever launching.
Started handoffd.bb manually with SWARMFORGE_PREVENT_SLEEP=0 — the script's own built-in escape hatch that skips the sleep-inhibitor wrapper entirely. Two stuck handoffs delivered within seconds.
cd /home/adamsl/swarmforge-runs/red-box-part4 && rm -f .swarmforge/daemon/stop
SWARMFORGE_PREVENT_SLEEP=0 nohup bb swarmforge/scripts/handoffd.bb "$(pwd)" \
> .swarmforge/daemon/handoffd.log 2>&1 &
disown
Tooling backlog item: the inhibitor wrapper should fail loud (log + exit non-zero) instead of swallowing the polkit error, or default SWARMFORGE_PREVENT_SLEEP=0 on WSL where interactive polkit is never available.
Each of the five non-specifier roles has its own .swarmforge/handoffs/{inbox,outbox,sent} inside its own worktree — there is no single shared handoff directory. The supervisor's first pass checked only the top-level path and reported two real, correctly-delivered handoffs as "not existing anywhere."
Corrected by direct find across every worktree; the supervisor applied the per-worktree-path rule for the rest of the run without recurrence.
Hardender's mutmut run --max-children 8 pushed the box into severe memory pressure (as low as 78Mi–376Mi free of 5.8GB, 0 swap). Linux's OOM killer doesn't scope to the offending cgroup here — it took down whatever process looked worst by its heuristic, which twice included unrelated live services.
Fixed by dropping to --max-children 2; verified dashboard-server and lettabot both recovered.
Tooling backlog item: heavy roles (mutation testing especially) should run under a memory-limited cgroup, or SwarmForge should document/enforce a --max-children ceiling based on available RAM rather than leaving it to each role's judgment.
A second, different mutmut failure — json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0) — hit during mutant generation, most likely cache corruption from the repeated interrupted runs in item 3. It had already reached 94.0% (2796/2975, 328 stable survivors) before failing. Accepted as good enough; not chased further.
After accepting 94% and explicitly instructing hardender to stop, it launched a new ~3019-mutant scoped run instead. That run was killed directly (kill -TERM / pkill -9). A second re-launch started essentially as the stop message was being delivered — almost certainly a command queued before the instruction was processed. The supervisor caught and killed that one too.
"Progressing steadily" is not the same as "complying with stop." The only reliable check is ps confirming the specific forbidden process is actually gone — a healthy-looking process tree can hide a role that quietly restarted the exact thing it was told to stop.
Hardender's pane reported "Still running (1696/3019, ~56%)" verbatim across six or seven separate capture-pane checks spanning roughly 15 minutes, while ps-visible worker processes kept advancing through different functions underneath. The real number — found by tailing the actual stdout task-output log directly — was 81% (2446/3019).
Rule going forward: never trust a role's self-printed percentage for a long-running process. Tail the real stdout/task-output log file — found under /tmp/claude-*/…/tasks/*.output for background subagents.
Told explicitly to write the final mutation numbers (94.0%, 2796/2975, 1378 killed, 328 survivors, ~737 timeouts) into its completion handoff, hardender acknowledged, correctly stopped, correctly committed, correctly sent a handoff — and the handoff body contained none of the numbers. This happened twice in a row after being asked again.
A role can genuinely comply with every procedural part of an instruction (stop, commit, send) while silently dropping the actual content requested. Verbal/pane acknowledgment is not proof content was written — the delivered file has to be read directly. On the second occurrence, EG chose to let it go since the numbers were already safe in the human-maintained memory file independently.
While processing a new batch of architect handoffs, hardender began investigating and preparing to run gherkin-mutator/acceptance-mutation tooling from .cache/acceptance-pipeline-specification — the exact same unverified external repo EG had already told cleaner to skip earlier over unverified-fetch concerns. Not literally mutmut, but the same spirit of "keep testing past the point EG called it done." The supervisor flagged it proactively as a possible violation and redirected hardender to just write up the final numbers — which is the same instruction that then surfaced incident 7's second occurrence.
A genuine structural bug, not a timing/patience issue: architect's done_with_current.sh wrap-up step only ever sent its completion handoffs to hardender, never to coder or cleaner. Root-caused by checking actual inbox files across every worktree (not pane text, not self-reports) and finding nothing addressed to coder/cleaner from architect at all.
Fixed with an explicit instruction to architect to relay via swarm_handoff.sh to both coder and cleaner, the same mechanism already used for hardender. Confirmed resolved by diffing git log across all worktrees over the following passes.
This exact pattern recurred at least three times with QA: a test run (property tests, then a verbose rerun, then again) would finish with no printed pass/fail summary, and QA would simply move on to the next step (acceptance/run.sh, etc.) as if it had passed. No failure signal was ever actually seen — but "no failure signal" and "confirmed pass" are different claims, and only the second is safe to act on.
Every time this was caught, the actual pytest summary line was findable — either in the pane text directly or by grepping a task-output log — and it did, in fact, pass. The gap is the habit, not (this time) a hidden failure.
After incident 7's pattern seemed to repeat with QA, direct investigation revealed something different: the git_handoff type is a hard 5-header format (type/to/priority/task/commit) with no free-text body field — not a role skipping the request, a genuine tool constraint. QA had correctly identified this itself when asked and offered two compliant options.
Tooling backlog item: this is worth fixing at the source — either add an optional free-text field to git_handoff, or make the type:note companion-handoff pattern (option 2 above) a documented, first-class convention instead of something each role has to improvise when it hits the wall.
Architect and hardender both hit the same Claude session-limit reset window (resets 10:10pm America/New_York) within the same run — architect showed a real stop-and-wait confirmation dialog, hardender showed a 99% context-usage warning shortly before. The supervisor's own background-subagent session was separately caught by the same limit mid-check, requiring a later resume. Not a SwarmForge defect per se, but worth tracking: if roles cluster their heavy work in the same time window, they'll tend to hit shared rate-limit walls together.
Several roles' panes showed unsubmitted "self-suggestion" text that looked like real queued input but was a no-op — Ctrl+U did not reliably clear it, and Enter did not reliably submit it. The only reliable check is a follow-up capture-pane confirming the text actually left the input box and something changed, not assuming a keystroke worked.
The "supervisor" babysitting the six roles is not part of SwarmForge itself — it's a background subagent spawned ad hoc, standing in for EG's approval role. Background subagents run until they run out of queued work, then exit; that's documented, correct behavior, not a crash. It resumes cleanly via a message to its agent id each time, but this run needed on the order of 40+ manual re-nudges across its lifetime, several times per approval cycle. Filed as a real usability gap even though it's not a defect: a recurring-cadence mechanism (a scheduled loop instead of ad hoc re-nudges) would remove the need for a human to notice and relay every blocked prompt individually.
Discovered live during this run's wind-down: a recurring cron job (fired every 2 minutes, re-enqueuing the exact same "nudge the supervisor" prompt) had been left running well after the run was wound down — it kept resurfacing the same message and made it look like a fresh request each time, when it was actually a stale scheduled job nobody had turned off.
Turn the checking cron ON when a SwarmForge run starts (so the supervisor gets nudged automatically without manual re-prompting — this is the fix for item 14 above). Turn it OFF the moment the run is wound down/converged — otherwise it keeps firing indefinitely and can be mistaken for a live, ongoing request. Check active jobs before assuming a repeated instruction is genuine.
| # | Issue | Class | Status |
|---|---|---|---|
| 1 | Handoff daemon silently dead (polkit/sleep-inhibit) | Infra | Fixed |
| 2 | Supervisor checked wrong handoff directory | Coordination | Fixed |
| 3 | Mutmut OOM killed production services | Infra | Fixed |
| 4 | Mutation cache JSON corruption | Infra | Accepted |
| 5 | Hardender ignored "stop" twice | Trust / compliance | Killed manually |
| 6 | Stale TUI progress percentage | Observability | Worked around |
| 7 | Requested docs silently dropped, twice | Trust / compliance | Tolerated |
| 8 | Hardender drifted into banned tangent | Trust / compliance | Redirected |
| 9 | Architect never notified coder/cleaner | Coordination (structural) | Fixed |
| 10 | "Moved on" treated as "passed" | Verification discipline | Caught each time |
| 11 | Handoff format has no free-text body | Tooling gap | Worked around |
| 12 | Simultaneous session-limit hit | Infra / scheduling | Observed |
| 13 | Ghost/placeholder TUI text | Observability | Recurring |
| 14 | One-shot supervisor needs re-nudging | Ergonomics | Expected, not fixed |
| 15 | Checking cron left on after wind-down | Ergonomics / process | Process reminder |
--max-children ceiling.git_handoff, or make the headers-plus-note-handoff pattern a documented convention.