ROL Finance — Red Box Project

Investigation report: the document-annotation "red box" system

Opened 2026-08-08 · this report 2026-08-09 · Owner: Claude Code, this session · Status: one symptom class fixed and shipped; verification pipeline had a real gap; two further defects found, unfixed

Headline finding

The narrow bug reported originally — "no red box is drawn at all" on check-style receipts — is fixed and confirmed live. But visually re-inspecting the two receipts used to verify that fix (expenses 2004 and 2006) on 2026-08-09 shows the new box is drawn in the wrong place on both: it encloses the check's payer/bank-header block, not the "Pay to the Order of … $…" line. Tracing why the automated pipeline (14/14 acceptance scenarios, property tests, unit tests, all passing) never caught this led to a bigger finding: every test added in this feature mocks the vision call — none of them ever sent a real image to the real model and checked where the box actually landed. Separately, a third and unrelated pre-existing defect was found live today on expense 1989 (an annual-giving-summary letter): the code's own "decisive score" confidence floor turns out not to be an absolute floor at all.

Status board

Fixed No box at all

Check-style receipts that previously produced no annotation now produce one. Confirmed live via /api/open-supporting-document on 2004 and 2006: highlighted:true.

New defect Wrong placement

The box that is drawn lands on the payer/bank-header block, not the payee+amount line, on both fixtures used to verify the fix. See Part 3.

Coverage gap Nothing tests real placement

Acceptance, property, and unit tests all mock the vision call. No test in the whole feature ever checks a real image against a real returned region.

Unrelated, pre-existing Two more bugs

#2005: receipt_url resolves to a file that doesn't exist (never in scope). #1989: _best_line's decisive-score gate lets a unique sub-threshold winner through — found live today, unrelated to this fix.

Part 1 — Background

EG's framing at the start of this: "Mazda needs work on her red-boxing skills." That framing was corrected at the time and holds after this deeper look too — red-boxing is dashboard/document_annotation.py, a fully separate dashboard feature. Mazda's code never touches it. The Mazda Trainer's own automated gate (dashboard/trainer/red-box-gate.ts) had already, independently and correctly, written "Do not coach Mazda… this is a dashboard defect" on every occurrence across four Trainer reports (2026-08-05 through 08-07, expenses 2004/2005/2006).

2026-08-06 → 08-07
Trainer reports repeatedly log "No high-confidence expense row was found in the image." for check-style/handwritten receipts (expenses 2004, 2006) and a separate available:false for expense 2005. Every report: red-box gate FAIL, explicitly not Mazda's fault.
2026-08-08, evening
Escalated to Suzuki's live bug-hunt pipeline twice. Both runs correctly triaged the root cause direction but returned a hallucinated, non-applicable patch — wrong function signatures/arity against the real code, confirming Suzuki's live-run mode has no execution access to verify a patch against real source.
2026-08-08, 21:32–22:43
Escalated further to a SwarmForge six-pack run (specifier → coder → cleaner → architect → hardener → QA) in an isolated scratch clone. Full pipeline detail in Part 2.
2026-08-08, 22:5x
5 commits merged into fix/intake-duplicate-rows on the live repo; dashboard restarted 2026-08-09 06:55 EDT and has been running the new code since.
2026-08-09, this session
EG reports the bug still exists. Live re-investigation (Part 3) confirms a real, more serious defect the shipped fix does not resolve, plus a third unrelated live bug.

Part 2 — What SwarmForge actually did

Run against an isolated scratch clone (~/swarmforge-runs/red-box-check-style-fallback, origin remote removed), never the live checkout. Every approval prompt across all six roles was reviewed and answered by hand over the course of the run; two attempts by architect and coder to reach into the live checkout's venv instead of building their own were caught and redirected. A real gap in this SwarmForge distribution — merge_and_process referenced in the constitution but not shipped in swarmforge/scripts/ — was found and patched with a one-line stopgap so every role could actually complete its handoff. Full operational writeup: mazda_suzuki_escalation_contract.md §4a.

Specifier
gpt-5.6-sol · high → low

14m31s of research (read the real document_annotation.py, the dashboard's highlight contract, cloned the official Acceptance-Pipeline-Specification tooling), then wrote two Gherkin feature files and two end-to-end QA suites — reviewed and approved before commit.

  • features/image-receipt-fallback-highlighting.feature — 3 scenario outlines: fallback succeeds on 2004/2006, fails closed on 4 distinct failure modes (no confident region / ambiguous / out-of-bounds / service unavailable), and an explicit regression scenario re-verifying two past-incident rows (1985 DTE duplicate-charge, 1522 APPLE.COM amount-column) are untouched.
  • features/image-receipt-fallback-strategy.feature — architectural invariants: fallback is called 0 times when the OCR path already matched, exactly 1 time when it didn't; _DECISIVE_SCORE stays literally 10; the IExpenseDocumentAnnotator contract (image/PDF/Excel selection) is unchanged.
  • qa/image-receipt-fallback-highlighting.md, qa/image-receipt-fallback-strategy.md — human-executable manual QA scripts naming the real archived files for 2004/2006 and asking a person to visually confirm box placement. These are the steps nobody, including me, actually ran by eye during the build — see Part 3.

9beea7d8 docs: specify image receipt fallback

Coder
gpt-5.6-sol · high → low

Built the new Strategy: IImageRegionFallbackMatcher (port) + CodexCliImageRegionFallbackMatcher (concrete adapter over the codex CLI already on this box, subprocess call, JSON-only response contract). Wired into ImageExpenseDocumentAnnotator.annotate so it only engages when _best_line already returned None, requires confidence ≥ 0.9, validates the returned region is within the image's real pixel bounds, and skips the OCR-specific row-expansion heuristics (which don't apply to a vision-returned region). The pre-existing OCR scoring path is untouched — _best_line/_line_score/_DECISIVE_SCORE are called exactly as before.

The prompt actually sent to the model (unchanged since coder wrote it):

Inspect only the attached supporting-document image. Find the single
compact source payment region for this expense:
date: {evidence.expense_date}
amount: {evidence.amount}
description/payee: {evidence.description}
vendor key: {evidence.vendor_key}
For a check image, select the check face containing the payment and
payee. Exclude endorsement, back-office, remote-deposit, posting,
summary, and unrelated regions. Coordinates must use original image
pixels. If there is no confident unique match, return no regions.
Return only JSON in this exact shape:
{"confidence":0.0,"regions":[{"left":0,"top":0,"right":0,"bottom":0}]}

3b736a74 feat(dashboard): add image receipt region fallback — 9 files, +786/−13

Cleaner
gpt-5.6-luna · high

Set up an isolated tmp/test-venv, ran the real acceptance pipeline (14/14 pass), ran CRAP/DRY checks against the Babashka acceptance tooling (clean, no duplicate candidates), then simplified the acceptance-generator's own control flow.

197c16eb refactor: simplify acceptance generator control flow — 2 files, +44/−28

Architect
gpt-5.6-sol · high → low

Split the single growing file into three: document_annotation.py (annotators), document_annotation_contracts.py (the abstract ports), codex_image_region_fallback.py (the new concrete adapter, isolated from the rest of the module). Added a dedicated architecture test and property tests. Ran the full dashboard suite in its own local venv: 632 passed, 2 skipped, 2 pre-existing failures correctly identified as unrelated.

95d8d82f refactor: enforce image fallback dependency boundaries — 8 files, +300/−197

Hardener
gpt-5.6-luna · medium

Own local venv, mutation-hardening tools (clj-mutate/crap4clj/dry4clj), full suite re-run (632 passed, 2 unrelated pre-existing failures), then hardened the acceptance-generator's own coverage.

bc2bbfa2 test: harden image receipt fallback acceptance coverage — 6 files, +155/−6

QA
gpt-5.6-luna · medium

Ran the full automated acceptance suite (acceptance/run.sh — the Babashka Gherkin runner) one final time: 14/14. Ran pyright (0 errors), CRAP/DRY (clean). Broadcast the terminal completion handoff. Did not — could not, from inside this pipeline — execute the human-facing manual QA scripts it had itself written (qa/image-receipt-fallback-highlighting.md's literal instructions: open the dashboard in a browser, click View Receipt, look at the image). That gap is the direct cause of Part 3.

No commit — no QA-owned changes were needed per its own verdict.

Part 2b — Reported verification results (as produced by the pipeline)

CheckResultWhat it actually verifies
Acceptance scenarios14 / 14 passedControl flow: fallback called the right number of times, highlighted flag set correctly, confidence threshold enforced. Not real image → real model → real pixels.
Property tests2 / 2 passedInvariants over the annotator contract, not image content.
Static types (pyright)0 errorsType correctness only.
CRAP / DRYcleanCode complexity/duplication metrics only.
Full dashboard suite632 passed, 2 skipped, 2 pre-existing failsConfirmed the 2 failures are unrelated (Claude 429 rate-limit tests, pre-existing before this work).

Every one of these is a legitimate, correctly-passing check for what it tests. None of them tests what a human means by "the red box is in the right place." That's the gap.

Part 3 — Live re-investigation, 2026-08-09

3.1 — The narrow original bug: confirmed fixed

Direct calls to the live dashboard API, today, against the exact two expenses from the original Trainer reports:

$ curl -sS -X POST http://localhost:8765/api/open-supporting-document \
    -H 'Content-Type: application/json' -d '{"expense_id": 2004, "document_type": "receipt"}'
{"ok": true, "url": "/supporting-document/2004/receipt", "highlighted": true, "highlight_note": ""}

$ curl -sS -X POST http://localhost:8765/api/open-supporting-document \
    -H 'Content-Type: application/json' -d '{"expense_id": 2006, "document_type": "receipt"}'
{"ok": true, "url": "/supporting-document/2006/receipt", "highlighted": true, "highlight_note": ""}

Both now return highlighted: true — the narrow "no box drawn at all" symptom is genuinely gone. No Trainer report since the 06:55 restart shows this failure recurring on a fresh scan either.

3.2 — But the box is in the wrong place, on both

Actually opening the two annotated images the fix produces — something no automated check in the pipeline does — shows the red box consistently landing on the check's payer/bank-header block at the top of the image, not on the "Pay to the Order of … $ …" line the prompt explicitly asked for.

Expense 2004 annotated receipt, box on the wrong region
Expense #2004 — River of Life Ministries check, $125.00 to John Roark, 2025-09-08. The box encloses the account-holder header ("RIVER OF LIFE MINISTRIES INC. / PASTORS R & R BARNES / …") — it contains no amount, no payee name, and no date.
Expense 2006 annotated receipt, box on the wrong region
Expense #2006 — Fifth Third check, $30.00 to Gabrielle McKay, 2025-05-02. The box encloses the bank security-notice bar and account/routing header (it does capture the correct $30.00 and 05/02/2025, coincidentally, but not the payee "GABRIELLE MCKAY" a few lines below).

3.3 — Why nothing caught this: every test mocks the vision call

Traced the acceptance-test step implementation and every relevant unit test. The pattern is the same everywhere: a fake stand-in returns a hand-written JSON string, so the test proves the plumbing around the call (does the annotator call the fallback the right number of times, does it respect the confidence floor, does it set the right flag) — never that a real model, given a real image, returns a geometrically correct answer.

acceptance/acceptance_steps.py — the step behind Gherkin's "exactly one red box encloses <target_region>":

def _assert_one_box(world, match, example):
    value = _require_choice(
        _example_value(example, match.group(1)),
        _TARGET_REGIONS.keys() | _CONFUSING_REGIONS,
        "receipt region",
    )
    if value in _TARGET_REGIONS and _TARGET_REGIONS[value] != world.get("expense"):
        raise AssertionError(...)
    _assert_annotated(world, match, example)      # only checks highlighted is True
    assert world["fallback"].calls <= 1           # only checks call count

It checks a symbolic label ("this scenario is about expense 2004's region") against a fixture dictionary — never the actual returned coordinates against the actual image content. The Gherkin's English sentence promises more than the step behind it verifies.

dashboard/tests/test_document_annotation.py — the fallback matcher's own "does it call Codex correctly" test:

def runner(command, **kwargs):
    calls.append((command, kwargs))
    class Completed:
        returncode = 0
        stdout = ('{"confidence":0.98,"regions":['
                   '{"left":100,"top":80,"right":900,"bottom":420}]}')
        stderr = ""
    return Completed()

match = CodexCliImageRegionFallbackMatcher(codex_path="/opt/codex", runner=runner)\
    .find_region("/receipts/check.jpg", target)

The region (100, 80, 900, 420) is simply invented by the test author — there is no /receipts/check.jpg, no real subprocess call, no real model response, anywhere in this test file or the property tests file. This is true of every single test added across all five commits.

3.4 — Confirming the raw mechanism itself does work

To separate "the CLI call is broken" from "the CLI call works but returns a geometrically imprecise answer," the exact subprocess call the shipped code makes was run standalone against both real fixture images, under an environment deliberately restricted to what the live systemd service actually provides (minimal PATH, no interactive-shell extras):

$ env -i HOME="$HOME" PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin" python3 -c '
    subprocess.run(["/usr/local/bin/codex","exec","--ephemeral","--ignore-user-config",
      "--ignore-rules","--skip-git-repo-check","--sandbox","read-only",
      "--model","gpt-5.4","--image", GRAND_RAPIDS_FIRST_JPG, "-"],
      input=PROMPT, capture_output=True, text=True, timeout=120)'

RETURNCODE: 0
STDOUT: {"confidence":0.94,"regions":[{"left":89,"top":86,"right":1138,"bottom":402}]}

Return code 0, clean JSON on stdout, confidence 0.94 — comfortably above the 0.9 floor. The mechanism is not crashing, timing out, or environment-broken. The model is confidently, successfully returning the wrong region — the top strip of the check, which does technically contain a date/account/routing summary, rather than the "Pay to the Order of" line the prompt asks for. This is a prompt/grounding quality problem, not an infrastructure problem — exactly the kind of thing only a real image, run for real, would ever surface, which is why the mocked test suite never saw it.

3.5 — A third, unrelated bug found live today (expense 1989)

This one is not part of the SwarmForge fix at all — for this document, _best_line's existing OCR path found a match (a real, non-None region), so the new fallback matcher was never even invoked. It is a pre-existing defect in code SwarmForge didn't touch, that happened to surface today.

Expense 1989 annotated document, box on the letterhead instead of the transaction table
Expense #1989 — Children's Vision Int. annual giving-summary letter, $3,047.00, 2026-01-06 (Window Scanner run, Trainer report 20260809-122318, which itself PASSED the red-box gate — the gate only checks "was something highlighted", the same gap as 3.3). The box encloses the vendor's own letterhead line ("Children's Vision Int. Inc" — repeated top-left/top-right) instead of the $3,047.00 total in the giving table below.

Ran the real OCR + scoring pipeline against this exact archived image, with the expense's real evidence (date 2026-01-06, amount 3047.00, description "Children's Vision Int. Inc."):

DECISIVE_SCORE = 10
Total OCR candidates: 359

  9  'Children's Vision Int. Inc'                      ← WINNER (the letterhead)
  9  'Children's Vision Int. Inc.'
  9  'Children's Vision Int. Inc. Children's Vision Int. Inc Carrera 26F No. 35-46 Sur'
  … (11 more letterhead-derived candidates, all scoring 9)
  5  '3,047.00 during 2025. This receipt is a detailed listing of your'   ← the real total line

WINNER region=(210, 178, 804, 216)  score=9  text='Children's Vision Int. Inc'

The winning score (9) is below _DECISIVE_SCORE (10) — by the constant's own name, this should not have been treated as decisive. Reading _best_line's actual gate (document_annotation.py, around the comment "For looser description matches, reject a tie instead of boxing the wrong repeated amount") shows why it won anyway:

if (
    ranked[0][0] < _DECISIVE_SCORE
    and rival is not None
    and rival[0] == ranked[0][0]
):
    return None, ranked[0][0], ""

This only rejects a sub-decisive score when a rival candidate ties it exactly. A unique top-scorer below the decisive threshold — no tie, nothing to reject against — falls through this check and wins by default. _DECISIVE_SCORE functions as a tie-breaking safety net, not the absolute confidence floor its name and surrounding comments imply. This document's letterhead repeats the vendor's own name twice in one visual line with no rival phrase scoring the same, so it wins uncontested at score 9 — one point under the bar meant to stop it.

Part 4 — What "bigger guns" should actually target

  1. Close the verification gap first, before touching the matching logic again. Add at least one real end-to-end test per document class (check, annual-summary letter, itemized receipt) that sends a real image to the real codex CLI (or a recorded real response, not a hand-typed one) and asserts the returned region's pixel bounds actually overlap the human-identified payee/amount area — not just that a region came back. Without this, any further fix is exactly as unverifiable as this one was.
  2. The vision prompt needs grounding work, not a bigger model. The raw mechanism works (0.94–0.98 confidence, valid JSON, correct amount/date recognized in one case) — it's choosing the wrong region of a multi-region document. Likely needs either a stricter prompt (explicitly rule out the payer/return-address block, not just "endorsement/back-office/posting"), a two-pass approach (find candidate regions, then pick the one containing recognizable payee text), or ground truth examples in the prompt.
  3. Fix the decisive-score gate as a real floor in _best_line (document_annotation.py) — a solo top-scorer below _DECISIVE_SCORE should never win just because nothing tied it. This is a small, well-isolated logic fix, independent of the SwarmForge work, and is now trivially reproducible (see 3.5's exact scoring dump).
  4. Someone needs to actually run the manual QA scripts SwarmForge's specifier wrote (qa/image-receipt-fallback-highlighting.md, qa/image-receipt-fallback-strategy.md) — open the dashboard in a browser, click View Receipt, look. They were written correctly; nobody executed them as a human before merge.
  5. Expense #2005 (receipt_url resolves to nothing — "The receipt document could not be found") remains untouched, as originally scoped. Likely cause, still unconfirmed: a date-misparse (year 2028 instead of 2026) upstream in rol_finances' parse_and_categorize.py. Separate investigation, different repo.

Appendix — commits, files, reproduction

CommitRoleMessageFiles
9beea7d8specifierdocs: specify image receipt fallback4
3b736a74coderfeat(dashboard): add image receipt region fallback9
197c16ebcleanerrefactor: simplify acceptance generator control flow2
95d8d82farchitectrefactor: enforce image fallback dependency boundaries8
bc2bbfa2hardenertest: harden image receipt fallback acceptance coverage6

Merged into fix/intake-duplicate-rows on the live repo, 2026-08-08 22:39 EDT. Dashboard restarted and running this code since 2026-08-09 06:55 EDT.

Reproduce 3.1/3.2curl -sS -X POST http://localhost:8765/api/open-supporting-document -H 'Content-Type: application/json' -d '{"expense_id": 2004, "document_type": "receipt"}'; then open /supporting-document/2004/receipt in a browser
Reproduce 3.5expense_id 1989 — /api/open-supporting-document {"expense_id":1989,"document_type":"receipt"}
Full writeup / memorynotes_plans_handoffs/mazda_suzuki_escalation_contract.md §4a