The narrow bug reported originally — "no red box is drawn at all" on check-style receipts — is fixed and confirmed live. But visually re-inspecting the two receipts used to verify that fix (expenses 2004 and 2006) on 2026-08-09 shows the new box is drawn in the wrong place on both: it encloses the check's payer/bank-header block, not the "Pay to the Order of … $…" line. Tracing why the automated pipeline (14/14 acceptance scenarios, property tests, unit tests, all passing) never caught this led to a bigger finding: every test added in this feature mocks the vision call — none of them ever sent a real image to the real model and checked where the box actually landed. Separately, a third and unrelated pre-existing defect was found live today on expense 1989 (an annual-giving-summary letter): the code's own "decisive score" confidence floor turns out not to be an absolute floor at all.
Check-style receipts that previously produced no annotation now produce one. Confirmed live via /api/open-supporting-document on 2004 and 2006: highlighted:true.
The box that is drawn lands on the payer/bank-header block, not the payee+amount line, on both fixtures used to verify the fix. See Part 3.
Acceptance, property, and unit tests all mock the vision call. No test in the whole feature ever checks a real image against a real returned region.
#2005: receipt_url resolves to a file that doesn't exist (never in scope). #1989: _best_line's decisive-score gate lets a unique sub-threshold winner through — found live today, unrelated to this fix.
EG's framing at the start of this: "Mazda needs work on her red-boxing skills." That framing was corrected at the time and holds after this deeper look too — red-boxing is dashboard/document_annotation.py, a fully separate dashboard feature. Mazda's code never touches it. The Mazda Trainer's own automated gate (dashboard/trainer/red-box-gate.ts) had already, independently and correctly, written "Do not coach Mazda… this is a dashboard defect" on every occurrence across four Trainer reports (2026-08-05 through 08-07, expenses 2004/2005/2006).
"No high-confidence expense row was found in the image." for check-style/handwritten receipts (expenses 2004, 2006) and a separate available:false for expense 2005. Every report: red-box gate FAIL, explicitly not Mazda's fault.fix/intake-duplicate-rows on the live repo; dashboard restarted 2026-08-09 06:55 EDT and has been running the new code since.Run against an isolated scratch clone (~/swarmforge-runs/red-box-check-style-fallback, origin remote removed), never the live checkout. Every approval prompt across all six roles was reviewed and answered by hand over the course of the run; two attempts by architect and coder to reach into the live checkout's venv instead of building their own were caught and redirected. A real gap in this SwarmForge distribution — merge_and_process referenced in the constitution but not shipped in swarmforge/scripts/ — was found and patched with a one-line stopgap so every role could actually complete its handoff. Full operational writeup: mazda_suzuki_escalation_contract.md §4a.
14m31s of research (read the real document_annotation.py, the dashboard's highlight contract, cloned the official Acceptance-Pipeline-Specification tooling), then wrote two Gherkin feature files and two end-to-end QA suites — reviewed and approved before commit.
features/image-receipt-fallback-highlighting.feature — 3 scenario outlines: fallback succeeds on 2004/2006, fails closed on 4 distinct failure modes (no confident region / ambiguous / out-of-bounds / service unavailable), and an explicit regression scenario re-verifying two past-incident rows (1985 DTE duplicate-charge, 1522 APPLE.COM amount-column) are untouched.features/image-receipt-fallback-strategy.feature — architectural invariants: fallback is called 0 times when the OCR path already matched, exactly 1 time when it didn't; _DECISIVE_SCORE stays literally 10; the IExpenseDocumentAnnotator contract (image/PDF/Excel selection) is unchanged.qa/image-receipt-fallback-highlighting.md, qa/image-receipt-fallback-strategy.md — human-executable manual QA scripts naming the real archived files for 2004/2006 and asking a person to visually confirm box placement. These are the steps nobody, including me, actually ran by eye during the build — see Part 3.9beea7d8 docs: specify image receipt fallback
Built the new Strategy: IImageRegionFallbackMatcher (port) + CodexCliImageRegionFallbackMatcher (concrete adapter over the codex CLI already on this box, subprocess call, JSON-only response contract). Wired into ImageExpenseDocumentAnnotator.annotate so it only engages when _best_line already returned None, requires confidence ≥ 0.9, validates the returned region is within the image's real pixel bounds, and skips the OCR-specific row-expansion heuristics (which don't apply to a vision-returned region). The pre-existing OCR scoring path is untouched — _best_line/_line_score/_DECISIVE_SCORE are called exactly as before.
The prompt actually sent to the model (unchanged since coder wrote it):
Inspect only the attached supporting-document image. Find the single
compact source payment region for this expense:
date: {evidence.expense_date}
amount: {evidence.amount}
description/payee: {evidence.description}
vendor key: {evidence.vendor_key}
For a check image, select the check face containing the payment and
payee. Exclude endorsement, back-office, remote-deposit, posting,
summary, and unrelated regions. Coordinates must use original image
pixels. If there is no confident unique match, return no regions.
Return only JSON in this exact shape:
{"confidence":0.0,"regions":[{"left":0,"top":0,"right":0,"bottom":0}]}
3b736a74 feat(dashboard): add image receipt region fallback — 9 files, +786/−13
Set up an isolated tmp/test-venv, ran the real acceptance pipeline (14/14 pass), ran CRAP/DRY checks against the Babashka acceptance tooling (clean, no duplicate candidates), then simplified the acceptance-generator's own control flow.
197c16eb refactor: simplify acceptance generator control flow — 2 files, +44/−28
Split the single growing file into three: document_annotation.py (annotators), document_annotation_contracts.py (the abstract ports), codex_image_region_fallback.py (the new concrete adapter, isolated from the rest of the module). Added a dedicated architecture test and property tests. Ran the full dashboard suite in its own local venv: 632 passed, 2 skipped, 2 pre-existing failures correctly identified as unrelated.
95d8d82f refactor: enforce image fallback dependency boundaries — 8 files, +300/−197
Own local venv, mutation-hardening tools (clj-mutate/crap4clj/dry4clj), full suite re-run (632 passed, 2 unrelated pre-existing failures), then hardened the acceptance-generator's own coverage.
bc2bbfa2 test: harden image receipt fallback acceptance coverage — 6 files, +155/−6
Ran the full automated acceptance suite (acceptance/run.sh — the Babashka Gherkin runner) one final time: 14/14. Ran pyright (0 errors), CRAP/DRY (clean). Broadcast the terminal completion handoff. Did not — could not, from inside this pipeline — execute the human-facing manual QA scripts it had itself written (qa/image-receipt-fallback-highlighting.md's literal instructions: open the dashboard in a browser, click View Receipt, look at the image). That gap is the direct cause of Part 3.
No commit — no QA-owned changes were needed per its own verdict.
| Check | Result | What it actually verifies |
|---|---|---|
| Acceptance scenarios | 14 / 14 passed | Control flow: fallback called the right number of times, highlighted flag set correctly, confidence threshold enforced. Not real image → real model → real pixels. |
| Property tests | 2 / 2 passed | Invariants over the annotator contract, not image content. |
| Static types (pyright) | 0 errors | Type correctness only. |
| CRAP / DRY | clean | Code complexity/duplication metrics only. |
| Full dashboard suite | 632 passed, 2 skipped, 2 pre-existing fails | Confirmed the 2 failures are unrelated (Claude 429 rate-limit tests, pre-existing before this work). |
Every one of these is a legitimate, correctly-passing check for what it tests. None of them tests what a human means by "the red box is in the right place." That's the gap.
Direct calls to the live dashboard API, today, against the exact two expenses from the original Trainer reports:
$ curl -sS -X POST http://localhost:8765/api/open-supporting-document \
-H 'Content-Type: application/json' -d '{"expense_id": 2004, "document_type": "receipt"}'
{"ok": true, "url": "/supporting-document/2004/receipt", "highlighted": true, "highlight_note": ""}
$ curl -sS -X POST http://localhost:8765/api/open-supporting-document \
-H 'Content-Type: application/json' -d '{"expense_id": 2006, "document_type": "receipt"}'
{"ok": true, "url": "/supporting-document/2006/receipt", "highlighted": true, "highlight_note": ""}
Both now return highlighted: true — the narrow "no box drawn at all" symptom is genuinely gone. No Trainer report since the 06:55 restart shows this failure recurring on a fresh scan either.
Actually opening the two annotated images the fix produces — something no automated check in the pipeline does — shows the red box consistently landing on the check's payer/bank-header block at the top of the image, not on the "Pay to the Order of … $ …" line the prompt explicitly asked for.
Traced the acceptance-test step implementation and every relevant unit test. The pattern is the same everywhere: a fake stand-in returns a hand-written JSON string, so the test proves the plumbing around the call (does the annotator call the fallback the right number of times, does it respect the confidence floor, does it set the right flag) — never that a real model, given a real image, returns a geometrically correct answer.
acceptance/acceptance_steps.py — the step behind Gherkin's "exactly one red box encloses <target_region>":
def _assert_one_box(world, match, example):
value = _require_choice(
_example_value(example, match.group(1)),
_TARGET_REGIONS.keys() | _CONFUSING_REGIONS,
"receipt region",
)
if value in _TARGET_REGIONS and _TARGET_REGIONS[value] != world.get("expense"):
raise AssertionError(...)
_assert_annotated(world, match, example) # only checks highlighted is True
assert world["fallback"].calls <= 1 # only checks call count
It checks a symbolic label ("this scenario is about expense 2004's region") against a fixture dictionary — never the actual returned coordinates against the actual image content. The Gherkin's English sentence promises more than the step behind it verifies.
dashboard/tests/test_document_annotation.py — the fallback matcher's own "does it call Codex correctly" test:
def runner(command, **kwargs):
calls.append((command, kwargs))
class Completed:
returncode = 0
stdout = ('{"confidence":0.98,"regions":['
'{"left":100,"top":80,"right":900,"bottom":420}]}')
stderr = ""
return Completed()
match = CodexCliImageRegionFallbackMatcher(codex_path="/opt/codex", runner=runner)\
.find_region("/receipts/check.jpg", target)
The region (100, 80, 900, 420) is simply invented by the test author — there is no /receipts/check.jpg, no real subprocess call, no real model response, anywhere in this test file or the property tests file. This is true of every single test added across all five commits.
To separate "the CLI call is broken" from "the CLI call works but returns a geometrically imprecise answer," the exact subprocess call the shipped code makes was run standalone against both real fixture images, under an environment deliberately restricted to what the live systemd service actually provides (minimal PATH, no interactive-shell extras):
$ env -i HOME="$HOME" PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin" python3 -c '
subprocess.run(["/usr/local/bin/codex","exec","--ephemeral","--ignore-user-config",
"--ignore-rules","--skip-git-repo-check","--sandbox","read-only",
"--model","gpt-5.4","--image", GRAND_RAPIDS_FIRST_JPG, "-"],
input=PROMPT, capture_output=True, text=True, timeout=120)'
RETURNCODE: 0
STDOUT: {"confidence":0.94,"regions":[{"left":89,"top":86,"right":1138,"bottom":402}]}
Return code 0, clean JSON on stdout, confidence 0.94 — comfortably above the 0.9 floor. The mechanism is not crashing, timing out, or environment-broken. The model is confidently, successfully returning the wrong region — the top strip of the check, which does technically contain a date/account/routing summary, rather than the "Pay to the Order of" line the prompt asks for. This is a prompt/grounding quality problem, not an infrastructure problem — exactly the kind of thing only a real image, run for real, would ever surface, which is why the mocked test suite never saw it.
This one is not part of the SwarmForge fix at all — for this document, _best_line's existing OCR path found a match (a real, non-None region), so the new fallback matcher was never even invoked. It is a pre-existing defect in code SwarmForge didn't touch, that happened to surface today.
20260809-122318, which itself PASSED the red-box gate — the gate only checks "was something highlighted", the same gap as 3.3). The box encloses the vendor's own letterhead line ("Children's Vision Int. Inc" — repeated top-left/top-right) instead of the $3,047.00 total in the giving table below.Ran the real OCR + scoring pipeline against this exact archived image, with the expense's real evidence (date 2026-01-06, amount 3047.00, description "Children's Vision Int. Inc."):
DECISIVE_SCORE = 10
Total OCR candidates: 359
9 'Children's Vision Int. Inc' ← WINNER (the letterhead)
9 'Children's Vision Int. Inc.'
9 'Children's Vision Int. Inc. Children's Vision Int. Inc Carrera 26F No. 35-46 Sur'
… (11 more letterhead-derived candidates, all scoring 9)
5 '3,047.00 during 2025. This receipt is a detailed listing of your' ← the real total line
WINNER region=(210, 178, 804, 216) score=9 text='Children's Vision Int. Inc'
The winning score (9) is below _DECISIVE_SCORE (10) — by the constant's own name, this should not have been treated as decisive. Reading _best_line's actual gate (document_annotation.py, around the comment "For looser description matches, reject a tie instead of boxing the wrong repeated amount") shows why it won anyway:
if (
ranked[0][0] < _DECISIVE_SCORE
and rival is not None
and rival[0] == ranked[0][0]
):
return None, ranked[0][0], ""
This only rejects a sub-decisive score when a rival candidate ties it exactly. A unique top-scorer below the decisive threshold — no tie, nothing to reject against — falls through this check and wins by default. _DECISIVE_SCORE functions as a tie-breaking safety net, not the absolute confidence floor its name and surrounding comments imply. This document's letterhead repeats the vendor's own name twice in one visual line with no rival phrase scoring the same, so it wins uncontested at score 9 — one point under the bar meant to stop it.
codex CLI (or a recorded real response, not a hand-typed one) and asserts the returned region's pixel bounds actually overlap the human-identified payee/amount area — not just that a region came back. Without this, any further fix is exactly as unverifiable as this one was._best_line (document_annotation.py) — a solo top-scorer below _DECISIVE_SCORE should never win just because nothing tied it. This is a small, well-isolated logic fix, independent of the SwarmForge work, and is now trivially reproducible (see 3.5's exact scoring dump).qa/image-receipt-fallback-highlighting.md, qa/image-receipt-fallback-strategy.md) — open the dashboard in a browser, click View Receipt, look. They were written correctly; nobody executed them as a human before merge.receipt_url resolves to nothing — "The receipt document could not be found") remains untouched, as originally scoped. Likely cause, still unconfirmed: a date-misparse (year 2028 instead of 2026) upstream in rol_finances' parse_and_categorize.py. Separate investigation, different repo.| Commit | Role | Message | Files |
|---|---|---|---|
| 9beea7d8 | specifier | docs: specify image receipt fallback | 4 |
| 3b736a74 | coder | feat(dashboard): add image receipt region fallback | 9 |
| 197c16eb | cleaner | refactor: simplify acceptance generator control flow | 2 |
| 95d8d82f | architect | refactor: enforce image fallback dependency boundaries | 8 |
| bc2bbfa2 | hardener | test: harden image receipt fallback acceptance coverage | 6 |
Merged into fix/intake-duplicate-rows on the live repo, 2026-08-08 22:39 EDT. Dashboard restarted and running this code since 2026-08-09 06:55 EDT.