ROL Finance · Evidence Model · Implementation Plan
Plan for Scanned Statements
A scanned bank statement is not the bank's own file, and it is not a receipt. It is a fourth, independent kind of evidence: proof that a paper copy was printed and reviewed. Today the intake pipeline has nowhere to put it, so it overwrites the real source document — or worse, invents a duplicate expense. This plan gives it its own slot, its own archive directory, its own View Scanned Statement button, and replaces five hardcoded three-way conditionals with one registry so a fifth kind is a one-line change.
Contents
- 1. The incident that forced this
- 2. The four evidence slots
- 3. Where files live
- 4. The design smell being removed
- 5. Interfaces
- 6. Fixing duplicate identity
- 7. The EVIDENCE_ATTACHED outcome
- 8. Dashboard & dialog
- 9. Migration & backfill
- 10. Failing tests to write first
- 11. Phase checklist
- 12. Decisions on record
1. The incident that forced this
On 2026-07-29 21:11 the Freezer scanner scanned a printed
Choice Privileges Mastercard 7580 statement covering
July 31 – August 15. Those same transactions had
already been in the books since 2026-07-02, loaded from the bank's own
yearly workbook choice_7580_year.xlsx. Here is what the system
did with them:
| Statement row | Amount | What the run did | What it should have done |
|---|---|---|---|
| PAYMENT — THANK YOU | credit | Skipped — correct | Skipped |
| KFC K980120 | $6.24 | Matched expense 1366, then tried to overwrite its
document_url → DocumentConflict |
Attach the scan to scanned_statement_url, leave
document_url alone |
| MR BURGER RESTAURANT 1 | $16.99 | Matched expense 1390, same overwrite attempt | Same — attach, don't overwrite |
| COUNTRY INN & SUITES — ELKHART | $179.08 | Not recognised. Queued as a brand-new expense. Row 1674 had already been created this way by the 07-27 Window scan, on top of existing row 1434 | Match row 1434 and attach the scan |
Then the whole batch rolled back on the KFC conflict, so the Country Inn row
appeared in neither expense_ids nor
duplicate_expense_ids. The scanner tab lists exactly the ids the
callback reports, so a $179.08 charge silently vanished from Verified
Transactions while sitting in the database twice.
document_url, which was already occupied by the
authoritative file. Everything downstream — the conflict, the rollback,
the duplicate row, the missing table row — follows from that one
missing slot.
supersedes_reference() in
e_two_e_processing/supporting_documents.py now treats a
transient incoming_scans/ path giving way to a durable archive
path as an upgrade rather than a conflict, at all three attach sites
in expense_repository.py (10 new tests, 39 passing). That stops
the rollback. It is a safety net, not the fix — under this plan a scan
never competes for document_url in the first place. Keep it:
staging paths are still transient and still must never win.
2. The four evidence slots
An expense row can be corroborated by up to four independent documents. Independent is the load-bearing word: they never substitute for each other, they accumulate. Each has its own column, its own button, its own archive root.
receipt_url
“View Receipt” — the merchant's own slip. Usually one
expense; lately sometimes several (itemised parents).
document_url
“View Source Document” — the bank's own file, as
downloaded: PDF or .xlsx. The authoritative record.
moms_ledger
“View Mom's Ledger” — Mom prints statements and
physically tapes them together. She has organised the checkbook this way
since the 1980s and never used a spreadsheet.
scanned_statement_url
“View Scanned Statement” — a scan of a
printed statement. Derived evidence: the paper implies a
downloadable original that may or may not be on file yet.
SOURCE slot that the primary
artifact belongs in. Marking the slot
is_derived_evidence = True is what lets policy code reason about
this without hardcoding the kind's name.
Provenance, not file type, picks the slot
A .jpg is not automatically a receipt and a statement is not
automatically SOURCE. The deciding input is how the
document arrived:
Classifier doc_type | Arrived via | Slot |
|---|---|---|
| statement / bank_statement / credit_card_statement | download (PDF, xlsx) | SOURCE |
| statement / bank_statement / credit_card_statement | scanner | SCANNED_STATEMENT |
| receipt / invoice | scanner or download | RECEIPT |
| moms_ledger | scanner | MOMS_LEDGER |
| other | any | none — halt, unsupported |
3. Where files live
Every database reference must name a durable, searchable
path — EG's attic rule: once the paper is filed away, the filename alone
has to be enough to find the scan. The scanner's
incoming_scans/scan_freezer_<ns>_<hash>.jpg fails both
halves: it is unsearchable, and it gets swept (a stray
git add -A destroyed two in-flight scans on 2026-07-29).
| Slot | Archive root | Layout |
|---|---|---|
SOURCE |
readable_documents/bank_statements/ |
{year}/{month}/{account}_{period}/ — unchanged |
RECEIPT |
readable_documents/receipts/ |
{year}/{month}/{day}/ — unchanged |
MOMS_LEDGER |
readable_documents/moms_ledger/ |
unchanged |
SCANNED_STATEMENT |
readable_documents/scanned_statements/ |
{year}/ — new; EG created
scanned_statements/2025/ |
Filename contract
Same naming discipline as the statement archive, so the file is findable by account and period rather than by timestamp:
readable_documents/scanned_statements/2025/
choice_privileges_mastercard_7580_july_31__august_15_scan.jpg
_supporting_document_roots() in dashboard/server.py
already whitelists all of ~/rol_finances/readable_documents, so
the new tree resolves and serves under the existing path-traversal check the
moment files appear in it. Do not add a root.
incoming_scans/ cleanup is a separate, later concern. A
move would break any in-flight intake still holding the staged path (that is
how the dashboard's red-box annotation locates the image it boxes).
4. The design smell being removed
Adding a fourth kind by copy-paste means editing every one of these, and missing one produces a button that renders but does not open, or a column that stores but never displays:
| File | Site | Shape today |
|---|---|---|
rol_finances supporting_documents.py | DocumentKind | 3 enum members + alias dict |
rol_finances expense_repository.py | attach, attach_many, consolidate_counterpart | allowed = {"receipt_url", "document_url", "moms_ledger"} ×3 |
rol_finances models.py | ExpenseRecord | 3 optional str fields |
dashboard server.py | _supporting_document_descriptors | hardcoded 3-tuple definitions |
dashboard server.py | open_supporting_document | field_by_type dict literal |
dashboard server.py | _supporting_document_path_for_expense | field_by_type dict literal (2nd copy) |
dashboard server.py | _supporting_document_view_for_expense | field_by_type dict literal (3rd copy) |
dashboard server.py | _resolve_local_supporting_document | if document_type == 'receipt' branch |
dashboard server.py | _matching_expense, lookup_supporting_documents, report queries | 3 field names spelled into ≥6 SELECT lists |
dashboard rol-finance-reports-controller.js | line ~484 | mk("button", "cp-view-receipt", "View Receipt") hardcoded |
5. Interfaces
Ports are abstract (ABC in dashboard code, Protocol
where rol_finances already uses structural typing). Concretes are
wired only in composition roots —
build_document_annotation_service(),
_get_supporting_document_annotation_service(),
js/dashboard-boot.js. Nothing below imports a concrete sibling.
5.1 The catalog — single source of truth
New module: rol_finances/e_two_e_processing/supporting_document_catalog.py
class ISupportingDocumentSlot(Protocol):
"""One kind of evidence an expense row can carry."""
kind: str # 'receipt' | 'source' | 'moms_ledger' | 'scanned_statement'
expense_field: str # DB column
label: str # dialog button text
archive_subdir: str # under readable_documents/
is_derived_evidence: bool # True when the artifact implies a primary elsewhere
class ISupportingDocumentCatalog(ABC):
"""The only place a supporting-document kind is declared."""
@abstractmethod
def slots(self) -> tuple[ISupportingDocumentSlot, ...]: ...
@abstractmethod
def slot_for_kind(self, kind: str) -> ISupportingDocumentSlot | None: ...
@abstractmethod
def slot_for_field(self, field: str) -> ISupportingDocumentSlot | None: ...
@abstractmethod
def fields(self) -> tuple[str, ...]:
"""Column names, in display order — for SELECT lists and INSERTs."""
Implementation DocumentKindCatalog derives itself
from the extended DocumentKind enum, so the enum stays the
declaration and the catalog stays the query surface. No second list to keep
in sync.
5.2 Classification routing
class Provenance(Enum):
DOWNLOADED = "downloaded" # pulled from the bank's site
SCANNED = "scanned" # came off a physical scanner
class IDocumentClassificationRouter(ABC):
"""(doc_type, provenance) -> the slot this artifact belongs in."""
@abstractmethod
def route(self, doc_type: str, provenance: Provenance) -> ISupportingDocumentSlot | None:
"""None means unsupported — the caller must halt, never guess."""
Implementation ProvenanceAwareRouter. This is the
one place that knows “a scanned statement is
SCANNED_STATEMENT, a downloaded one is SOURCE”.
Fails closed, matching router/classify.py's existing rule that a
wrong guess is worse than no guess.
5.3 Evidence upgrade policy
Generalises the supersedes_reference() function already landed
into an injected policy, so the repository stops hardcoding it:
class IEvidenceUpgradePolicy(ABC):
@abstractmethod
def supersedes(self, existing: str, proposed: str) -> bool:
"""True when `proposed` is strictly better evidence for the same slot,
so replacing it is an upgrade rather than a human-resolvable conflict."""
- StagingToArchiveUpgradePolicy — the rule
shipped 2026-07-29: transient
incoming_scans/→ durable archive is an upgrade; the reverse never is; two durable references still conflict. - NeverUpgradePolicy — strict fail-closed, for tests and for slots where any replacement needs a human.
Inject into ExpenseRepository (default
StagingToArchiveUpgradePolicy) and replace the three direct
supersedes_reference(...) calls with
self.upgrade_policy.supersedes(...).
5.4 Archiving
@dataclass(frozen=True)
class ArchivedDocument:
path: str # durable absolute path — what goes in the DB
slot_kind: str
copied: bool # False when an identical file was already archived
class IDocumentArchiveLocator(ABC):
"""Where a document of this slot, for this date, belongs."""
@abstractmethod
def archive_dir(self, slot: ISupportingDocumentSlot, when: date) -> Path: ...
@abstractmethod
def archive_filename(self, slot: ISupportingDocumentSlot, meta: Mapping) -> str:
"""Searchable name: account + period, never a timestamp or hash."""
class IDocumentArchiver(ABC):
@abstractmethod
def archive(self, staging_path: str, slot: ISupportingDocumentSlot,
meta: Mapping) -> ArchivedDocument:
"""COPY the file to its durable home. Idempotent by content hash."""
- BankStatementArchiveLocator — wraps the
existing
{year}/{month}/logic instatement_archive.py; no behaviour change. - ScannedStatementArchiveLocator — new,
scanned_statements/{year}/. - ReceiptArchiveLocator — wraps the
existing
{year}/{month}/{day}/receipt layout. - CopyOnWriteArchiver — the single
IDocumentArchiver, given a locator. Copies, hashes, returnscopied=Falseon a byte-identical re-scan.
5.5 Attaching evidence to an existing row
class IEvidenceAttachmentService(ABC):
"""Fortify an already-stored expense with additional evidence.
Distinct from SupportingDocumentService.attach: that answers 'put this
reference in that field'. This answers the intake question — 'this document
proves a row we already have; record that and store nothing new'.
"""
@abstractmethod
def attach_evidence(self, expense_ids: Sequence[int],
slot: ISupportingDocumentSlot,
archived: ArchivedDocument) -> EvidenceAttachmentReport: ...
@dataclass(frozen=True)
class EvidenceAttachmentReport:
attached_expense_ids: tuple[int, ...]
already_attached_ids: tuple[int, ...]
conflicted_ids: tuple[int, ...] # needs a human — never overwritten
slot_kind: str
5.6 Dashboard resolution & presentation
New module: dashboard/supporting_document_slots.py
class ISupportingDocumentResolver(ABC):
"""A stored reference -> a local file the viewer may serve, or None."""
@abstractmethod
def resolve(self, reference: str, slot: ISupportingDocumentSlot) -> str | None: ...
class IDocumentReferenceFallback(ABC):
"""What else we know when the stored reference no longer resolves.
Currently only 'source' has one (_report_source_document_reference).
SCANNED_STATEMENT needs its own: the intake's staged scan image.
"""
@abstractmethod
def fallback_reference(self, expense_row: Mapping, report_path: str) -> str: ...
class ISupportingDocumentDescriptorFactory(ABC):
"""The dialog's button list. Driven by the catalog, not a literal tuple."""
@abstractmethod
def descriptors(self, expense_row: Mapping, report_path: str = "") -> list[dict]: ...
- ReceiptIndexResolver — the existing
_resolve_receipt_url_pathfilesystem index. - AllowedRootsResolver — the existing generic path-traversal-checked resolution.
- ChainedResolver — tries an ordered list;
replaces the
if document_type == 'receipt'special case with per-slot configuration. - IntakeScanFallback — for
SCANNED_STATEMENT: the intake pointer's immutableimage_path. Must never fall back to the scanner's reusablescan_freezer.jpg, which by now holds a different document. - CatalogDescriptorFactory — loops the catalog; each slot may carry a fallback.
IExpenseDocumentAnnotationService and
IExpenseDocumentAnnotator in
dashboard/document_annotation.py already dispatch by file type,
and a scanned statement is an image —
ImageExpenseDocumentAnnotator handles it as-is, including the
widened red box for rows whose amount EG's pen made unreadable. Register the
new kind in build_document_annotation_service() and stop.
5.7 Front end
New: js/abstract/supporting-document-slots.interface.js
// Pure: no DOM, no fetch. Unit-tested under js/tests/.
export const SLOT_ORDER = ['receipt', 'source', 'scanned_statement', 'moms_ledger'];
export class SupportingDocumentButtonPolicy {
/** descriptors (from the server) -> ordered button specs.
* Unknown kinds are dropped, not guessed: an older server that has never
* heard of scanned_statement must still render the other three. */
buttonsFor(descriptors) { /* ... */ }
}
js/implementation/rol-finance-reports-controller.js stops
constructing named buttons and renders whatever the policy returns. After
this, adding a fifth kind requires no JS change at all.
6. Fixing duplicate identity
Row 1674 should never have existed. It exists because duplicate detection keys on description wording, and the same charge is worded three different ways by three different sources:
| Source | Description | Derived vendor key |
|---|---|---|
| xlsx (row 1434) | COUNTRY INN & SUITES - ELKHART ,IN | country_inn_by_carlson |
| 07-27 scan (row 1674) | COUNTRY INN & SUITES - ELKHART IN ARRIVE 08/14/25 DEPART 08/15/25 FOLIO#0834252978 | country_inn_suites_elkhart_in_arrive_… |
| 07-29 scan | (OCR variant again) | — matched neither |
DuplicateChecker.is_duplicate matches on exact
id_light or exact (date, amount, description).
find_fuzzy_duplicate then requires one vendor key to be a
prefix of the other — country_inn_by_carlson and
country_inn_suites_elkhart… share the leading token
country_inn but neither prefixes the other, so it reports no
match.
class IDuplicateIdentityStrategy(ABC):
"""One way of asking 'is this printed transaction already in the books?'"""
@abstractmethod
def find_existing(self, candidate: TransactionIdentity) -> DuplicateVerdict: ...
@dataclass(frozen=True)
class DuplicateVerdict:
expense_ids: tuple[int, ...] # 0 = new, 1 = duplicate, >1 = AMBIGUOUS
strategy: str # which rule fired — goes in the trace
decisive: bool # may this outrank a weaker strategy?
Chain, strongest first, in ChainedDuplicateIdentity:
- ExactIdLightIdentity — today's primary check.
- DateAmountDescriptionIdentity — today's secondary.
- DateAmountPayeeHeadIdentity — new.
Same date, same amount, and a shared leading payee token run. Reuses the
scoring already proven in
document_annotation.py's_payee_head_tokens, which exists for exactly this problem on the OCR side. - DateAmountSoleCandidateIdentity —
new, non-decisive. Same date, same amount, exactly one candidate
in the whole table → duplicate. Two or more →
AMBIGUOUS.
len(expense_ids) > 1 is never resolved by picking one. It
quarantines to _needs_review/ and raises a statement-review
dialog, exactly like an unreadable amount does today. Guessing wrong attaches
a document to the wrong charge, which is worse than asking.
AMBIGUOUS for that pair rather than
a duplicate. The merge is a prerequisite for a clean re-scan, not an
afterthought.
7. The EVIDENCE_ATTACHED outcome
Today a statement run reports stored and duplicates.
A re-scan that attaches new evidence to known rows is neither — it is a
successful run that stored nothing and improved the books. It needs
its own name, or the judge keeps failing a correct outcome (which it did on
2026-07-29: verdict FAIL on statement_transactions_stored for a
run that behaved properly).
{
"ok": true,
"outcome": "EVIDENCE_ATTACHED",
"transactions_parsed": 4,
"skipped_credits": 1,
"duplicates": 3,
"duplicate_expense_ids": [1366, 1390, 1434],
"stored": 0,
"expense_ids": [],
"scanned_statement_attached": [1366, 1390, 1434],
"scanned_statement_path": ".../scanned_statements/2025/choice_..._scan.jpg",
"conflicted_expense_ids": []
}
Consequences to implement:
- STEP 8 template (
build_mazda_scan_message()indashboard/server.py) must document the new field and the new outcome, and must state that a scanned statement attaches toscanned_statement_url— neverdocument_url. The existing STEP 8 wording that maps “statement →document_url” is now wrong for scans and is exactly the instruction that produced this incident. - The intake judge rubric
(
rol_finances/tools/self_improving_agent, served bymazda-tools-mcp.serviceon:8791) must acceptEVIDENCE_ATTACHEDas a pass. Restart that unit after editing the rubric. _fold_event_into_intakemust unionscanned_statement_attachedinto the intake's displayed ids, so an attach-only run still renders its rows. This is the direct fix for “the Country Inn row is not even showing”.- Rolled-back rows must be reported. Add
rolled_back_row_count. A run that discards rows while reporting only the survivors is how a $179.08 charge disappeared without a trace; the page must be able to say “3 parsed, 2 shown, 1 held back”.
8. Dashboard & dialog
Set Category dialog
The dialog gains a fourth button. Buttons render only when the document
actually resolves and is a viewable type — the existing
available flag rule is unchanged, just computed per catalog slot
instead of per hardcoded tuple entry.
[ View Receipt ] [ View Source Document ] [ View Scanned Statement ] [ View Mom's Ledger ]
^^^^^^^^^^^^^^^^^^^^^^^^^ new
Server endpoints — no new routes, all four kinds flow through the existing ones:
| Endpoint | Change |
|---|---|
POST /api/supporting-documents |
Returns a 4th descriptor; adds scanned_statement_url to the payload |
POST /api/open-supporting-document |
document_type: 'scanned_statement' resolves via the catalog |
GET /supporting-document/<id>/scanned_statement |
Works automatically once the catalog drives the type check |
Verified Transactions
Rows already carry a receipt marker. A scanned-statement marker is the same
pattern — batch-probed like /api/receipts-present so the
table does not fire one request per row. Distinct glyph/corner from the
receipt marker; the point is telling at a glance which rows are backed by
paper EG has physically reviewed.
Project Plans tab
Two edits in dashboard/dashboard.html, matching the existing plan tabs:
<!-- sub-nav, beside "Mazda Dev Status" -->
<button type="button" class="tab" data-nav="plans"
data-target="plans-scanned-statements">Scanned Statements</button>
<!-- view -->
<section id="plans-scanned-statements" class="view">
<iframe id="scanned-statements-plan-frame" class="plan-frame"
src="/notes_plans_handoffs/scanned_statements_plan.html"></iframe>
</section>
server.py already serves static files from
REPO_ROOT, so no route is needed. Verify with:
curl -s -o /dev/null -w '%{http_code}\n' \
http://localhost:8765/notes_plans_handoffs/scanned_statements_plan.html
9. Migration & backfill
9.1 Schema
Next in sequence after 2026_07_29_004_report_categories.py, both
directions, per the existing convention:
-- migrations/2026_07_30_005_expenses_add_scanned_statement_url.sql
ALTER TABLE expenses
ADD COLUMN scanned_statement_url VARCHAR(1024) NULL AFTER moms_ledger;
-- migrations/2026_07_30_005_expenses_add_scanned_statement_url_down.sql
ALTER TABLE expenses DROP COLUMN scanned_statement_url;
Nullable, no default, no backfill in the migration itself — the column
going in must not change a single existing row's meaning. Add the matching
optional field to ExpenseRecord in
e_two_e_processing/models.py and to the
_insert_with_cursor column list.
9.2 Merge the duplicated Country Inn charge
choice_7580_year.xlsx, which is precisely what the
SOURCE slot is for. Row 1674 is the row that, in EG's words,
“should never have been created”. Both already carry
category_id 160, so nothing is lost in the merge.
| Row | Fate | Ends up with |
|---|---|---|
| 1434 — from the xlsx | KEEP |
document_url → …/january/choice_7580_year/choice_7580_year.xlsxscanned_statement_url → …/scanned_statements/2025/choice_…_scan.jpg
|
| 1674 — from the 07-27 scan | DELETE after folding its evidence into 1434 | — |
Use the primitive that already exists rather than hand-written SQL:
SupportingDocumentService.consolidate_counterpart(canonical_id=1434,
duplicate_id=1674) atomically folds references and refuses if the
duplicate owns receipt metadata. Then delete 1674.
9.3 Backfill: scans currently sitting in document_url
Rows 1366, 1390, 1674 and roughly 64 legacy rows hold a scan path — many
of them a transient incoming_scans/ path — in
document_url, the slot that belongs to the bank's own file. Each
needs its reference moved, not copied:
# For each row whose document_url names a scan rather than a bank download:
service.reclassify(
expense_id=1366,
source_kind=DocumentKind.SOURCE,
target_kind=DocumentKind.SCANNED_STATEMENT,
reference=archived_scan_path,
)
# document_url is then free for the real downloaded statement.
reclassify() / move_reference() already exist and are
already atomic — this is the operation they were built for. Write the
backfill as a dry-run-first script with a JSON journal under
backups/, matching
document_url_archive_backfill_<stamp>.json.
incoming_scans/ or
scanned_statements/, or an image extension
(.jpg/.png) where the archive sibling is a
.pdf/.xlsx. Rows that are ambiguous go on a report
for EG rather than being moved on a guess.
10. Failing tests to write first
Stubs, grouped by suite. Each should fail for the right reason before any implementation lands.
rol_finances/tests/test_supporting_document_catalog.py
test_catalog_exposes_four_slots
test_scanned_statement_slot_maps_to_scanned_statement_url
test_scanned_statement_is_marked_derived_evidence
test_source_slot_is_not_derived_evidence
test_slot_for_field_round_trips_every_slot
test_fields_order_is_stable_for_select_lists
test_unknown_kind_returns_none_not_a_default # fail closed
test_document_kind_enum_and_catalog_cannot_drift
rol_finances/tests/test_document_classification_router.py
test_downloaded_statement_routes_to_source
test_scanned_statement_routes_to_scanned_statement
test_scanned_receipt_routes_to_receipt
test_downloaded_receipt_routes_to_receipt
test_moms_ledger_routes_to_moms_ledger
test_doc_type_other_routes_to_none
test_unknown_doc_type_routes_to_none_never_guesses
test_scanned_statement_never_routes_to_source # the incident, pinned
rol_finances/tests/test_evidence_upgrade_policy.py
test_archive_supersedes_staging
test_staging_never_supersedes_archive
test_two_durable_references_conflict
test_identical_reference_is_not_an_upgrade
test_never_upgrade_policy_rejects_everything
test_repository_uses_injected_policy_not_module_function
rol_finances/tests/test_document_archiver.py
test_scanned_statement_lands_under_scanned_statements_year
test_bank_statement_still_lands_under_bank_statements_year_month
test_archive_filename_is_account_and_period_not_timestamp
test_archive_copies_and_leaves_staging_file_in_place
test_reArchiving_identical_bytes_reports_copied_false
test_archive_never_writes_outside_readable_documents
rol_finances/tests/test_duplicate_identity.py
test_exact_id_light_matches
test_date_amount_description_matches
test_payee_head_matches_across_reworded_descriptions # country_inn_by_carlson
# vs country_inn_suites_elkhart
test_sole_date_amount_candidate_matches
test_two_candidates_same_date_amount_is_ambiguous_not_a_match
test_ambiguous_verdict_quarantines_and_does_not_insert
test_decisive_strategy_outranks_non_decisive
test_no_candidate_reports_new_expense
rol_finances/tests/test_evidence_attachment_service.py
test_attach_evidence_fills_empty_slot
test_attach_evidence_is_idempotent_for_same_reference
test_attach_evidence_never_overwrites_a_durable_reference
test_conflicted_ids_are_reported_not_raised
test_attaching_scanned_statement_leaves_document_url_untouched # the incident
test_attach_evidence_stores_no_new_expense_rows
rol_finances/tests/test_store_statement_evidence_outcome.py
test_rescan_of_stored_statement_reports_evidence_attached
test_evidence_attached_run_reports_ok_true
test_evidence_attached_lists_every_touched_expense_id
test_rolled_back_rows_are_counted_and_reported
test_partial_conflict_does_not_discard_clean_rows # the $179.08 loss
test_credit_rows_still_skipped
dashboard/tests/test_supporting_document_dialog.py (extend)
test_dialog_offers_four_buttons_when_all_documents_present
test_scanned_statement_button_label_is_view_scanned_statement
test_scanned_statement_button_hidden_when_field_empty
test_scanned_statement_button_hidden_when_file_missing_from_disk
test_button_order_is_receipt_source_scanned_statement_moms_ledger
test_descriptors_come_from_catalog_not_a_hardcoded_tuple
test_unknown_slot_kind_is_dropped_not_rendered
dashboard/tests/test_scanned_statement_view.py (new)
test_open_scanned_statement_resolves_archive_path
test_open_scanned_statement_falls_back_to_intake_staged_image
test_fallback_never_uses_reusable_scan_freezer_jpg
test_scanned_statement_is_annotated_by_image_annotator
test_scanned_statement_outside_readable_documents_is_refused
test_viewer_url_is_expense_scoped_and_stable
dashboard/tests/test_server.py (extend)
test_step8_template_routes_scanned_statement_to_its_own_field
test_step8_template_forbids_statement_scan_in_document_url
test_fold_event_unions_scanned_statement_attached_ids
test_evidence_attached_intake_renders_verified_transactions
test_intake_reports_rolled_back_row_count
dashboard/js/tests/supporting-document-slots.test.js (new)
buttonsFor returns descriptors in SLOT_ORDER
buttonsFor omits unavailable descriptors
buttonsFor drops unknown kinds instead of guessing a label
buttonsFor tolerates a server that omits scanned_statement entirely
buttonsFor returns an empty list for an expense with no evidence
.venv/bin/python -m pytest tests/ in each
repo; JS with bun test js/tests. The dashboard's
conftest.py autouse fixture already disables the Trainer and
redirects recent_report.json, so new intake tests inherit that.
11. Phase checklist
| Step | Lands in | |
|---|---|---|
| Phase 0 — already done | ||
| ✓ | supersedes_reference() + 3 repository call sites; 10 tests | rol_finances |
| ✓ | scanned_statements/2025/ directory created | rol_finances |
| ✓ | This plan, linked from Project Plans | letta-code |
| Phase 1 — declare the kind (no behaviour change) | ||
| Failing tests for catalog + router | rol_finances | |
DocumentKind.SCANNED_STATEMENT + aliases | supporting_documents.py | |
ISupportingDocumentCatalog + DocumentKindCatalog | new module | |
Migration 005 + ExpenseRecord field | rol_finances | |
| Phase 2 — make the catalog authoritative | ||
Repository allowed sets read catalog.fields() (3 sites) | expense_repository.py | |
IEvidenceUpgradePolicy injected into the repository | rol_finances | |
Dashboard: delete all 3 field_by_type literals | server.py | |
ISupportingDocumentResolver chain replaces the receipt if | server.py | |
CatalogDescriptorFactory replaces the definitions tuple | server.py | |
SELECT lists built from catalog.fields() | server.py | |
| Phase 3 — archiving & attachment | ||
IDocumentArchiveLocator + 3 locators; CopyOnWriteArchiver | rol_finances | |
IEvidenceAttachmentService | rol_finances | |
statement_archive.py delegates to the locators | rol_finances | |
| Phase 4 — duplicate identity | ||
IDuplicateIdentityStrategy + 4 strategies + chain | rol_finances | |
DuplicateChecker delegates to the chain | rol_finances | |
Ambiguous verdict quarantines to _needs_review/ | rol_finances | |
| Phase 5 — the pipeline outcome | ||
EVIDENCE_ATTACHED in the store's report | store_statement_transactions.py | |
| Clean rows survive a conflict on a different row | rol_finances | |
| STEP 8 template rewritten for the 4-slot model | server.py | |
Judge rubric accepts EVIDENCE_ATTACHED; restart mazda-tools-mcp | rol_finances | |
_fold_event_into_intake unions attached ids + rolled_back_row_count | server.py | |
| Phase 6 — the button | ||
supporting-document-slots.interface.js + bun tests | js/abstract/ | |
| Dialog renders from descriptors, no named buttons | js/implementation/ | |
| Scanned-statement row marker (batch-probed) | dashboard | |
| Injector re-run across all reports | rol_finances | |
| Phase 7 — data (needs EG's go-ahead) | ||
| Merge 1674 into 1434, delete 1674 | MySQL | |
Backfill ~64 scan references out of document_url | MySQL | |
| Re-scan the Choice statement to prove the whole path | Freezer | |
12. Decisions on record
- New column
scanned_statement_url VARCHAR(1024) NULL, aftermoms_ledger.- New kind
SCANNED_STATEMENT, classification stringscanned_statement.- New button
- “View Scanned Statement”, third of four, before Mom's Ledger.
- New archive root
readable_documents/scanned_statements/{year}/. Downloads stay inbank_statements/.- Never overwrite
- A scan is additive evidence. It fills its own slot or reports a conflict; it never displaces the bank's own file.
- Fail closed
- Ambiguous duplicate, unroutable
doc_type, unresolvable reference — all halt for a human. A wrong attachment is worse than a pause. - Keeper row
- 1434 (from the bank's xlsx), not 1674 (from the scan). Reverses the earlier call — see §9.2.
- One registry
- A fifth kind must cost one catalog entry and zero JS changes. If it costs more, Phase 2 is not finished.