Dashboard Architecture · Living Progress Plan

Server Rewrite

Incrementally reduce dashboard/server.py to a composition root. Every side effect crosses an injected Interface; every behavior lives in a focused Object or pure function; every extraction is pinned by tests before it moves.

Created 2026-07-29 Owner: dashboard teams Method: strangler refactor No route-contract changes

Progress

0 / 55
Interface/Object slices extracted
55 active rows remain in this document.
10,486server.py lines at baseline
10,962current audit 2026-07-31; +476 vs baseline
269top-level functions
233top-level assignments
1,062DashboardHandler lines
5,625test_server.py lines
532 pass7 skip; 6 dependency failures
First gate: make the baseline suite green. The six failures observed on 2026-07-29 are missing PyMuPDF, Pillow, and openpyxl in the dashboard venv. Refactoring does not begin until failures reliably mean regressions.

Operating rules

  1. Program to the Interface. Domain services import ports, never concrete adapters.
  2. Wire once. Concrete construction belongs in server.py or a dedicated application factory.
  3. TDD is mandatory. No production Interface or Object is written before its failing test. Work proceeds Red → Green → Refactor.
  4. Characterize before moving. Characterization tests pin public output, errors, caching, locking, and failure behavior before legacy code moves.
  5. Types are executable contracts. Python boundaries use strict Pydantic models; new browser code uses strict TypeScript.
  6. Keep pure code simple. Pure transforms remain functions; do not create an Interface without a boundary.
  7. Side effects require ports. Network, process, filesystem, database, clock, logging, metrics, and notifications are injected.
  8. One slice per change. Interface, Object, adapter, tests, wiring, and compatibility shim travel together.
  9. No live dependencies in unit tests. Use small fakes that implement the same port contract.
  10. Preserve HTTP contracts. Status, headers, JSON shapes, timeouts, and route names stay stable unless separately approved.
  11. Measure every slice. Record server.py lines immediately before and after; explain every non-negative delta and distinguish concurrent edits from slice-attributable movement.
  12. Delete completed detail. Once a row meets its exit test, remove it from this ledger. Git is the completion history.
The plan intentionally gets shorter. Do not grow a permanent “Completed” table. Update the baseline total only when a genuinely new boundary is discovered; otherwise remove finished rows and let the progress calculation advance.

Strict typed contracts and test-first development

Non-negotiable: Pydantic v2 is mandatory for Python data crossing an Interface, process, persistence, or HTTP boundary. All new frontend code is TypeScript. The type checker and the tests are merge gates, not advisory tools.

Python contract policy

ABCs describe behavior; Pydantic models describe data. Commands, results, API requests/responses, repository records, configuration, events, and adapter payloads use named models derived from one strict base:

from pydantic import BaseModel, ConfigDict

class StrictModel(BaseModel):
    model_config = ConfigDict(strict=True, extra="forbid", frozen=True)

Static checking runs with pyright --project dashboard/pyrightconfig.json, where typeCheckingMode = "strict". Runtime Pydantic validation complements static checking; neither replaces the other.

TypeScript contract policy

TDD order for every ledger row

  1. RED: write the consumer behavior test and run it. Record the expected failure; import/contract absence counts as red for a new extraction.
  2. Add Characterization tests around the legacy seam so the behavior being preserved is explicit.
  3. Add Shared contract tests that every concrete adapter and fake must pass.
  4. Add Negative validation tests proving malformed types, missing fields, extra fields, and invalid variants are rejected.
  5. GREEN: add only enough typed Interface/Object code to make the focused tests pass.
  6. REFACTOR: remove duplication and improve names while the tests and type checkers stay green.
  7. Run the focused suite, complete Python suite, complete browser suite, static type checks, and coverage gate.

Required test layers

LayerPurposeRequired evidence
Characterization testsFreeze current route/service behavior before movementSuccess, error, timeout, cache, and concurrency cases
Domain unit testsDrive service behavior through fake portsNo network, filesystem, process, clock, or database access
Shared contract testsHold real adapters and fakes to one behavioral contractParameterized suite runs against every implementation
Negative validation testsProve strict Pydantic/runtime TypeScript validationWrong, missing, extra, and ambiguous data fails closed
Adapter integration testsVerify serialization and external protocol mappingFixture servers/files/processes only; no production systems
Composition testsVerify every port is wired once to a valid concreteApplication builds and all routes register without serving
HTTP/browser testsProtect public routes, headers, JSON, navigation, and renderingBlack-box request and DOM assertions
New or extracted domain/application modules require 100% branch coverage. Coverage does not prove test quality: the test must also fail when the implemented behavior is removed or inverted. Adapter exceptions must be narrow and recorded in the decision log.

Target architecture

dashboard/
  server.py                 # construct adapters, start threads/server, compatibility exports
  http/                     # route registry, controllers, HTTP/WebSocket adapters
  agents/                   # Letta gateway, directory, messaging, agent health
  server_management/        # health probes, logs, SSH, restart orchestration
  model_usage/              # usage sources, history, provider health/failover
  intake/                   # scanners, classification, staging, dispatch, Trainer
  finance/                  # expense/receipt/report repositories and services
  voice/                    # existing transcription/cleanup pipeline
  tests/                    # mirrors the production domains above

The directory names are targets, not a mandate to move everything at once. A slice may begin as one flat module when that keeps the change reviewable.

Definition of done for the rewrite

Active Interface/Object ledger

Each work row is one reviewable extraction. The checkbox is deliberately visual rather than browser-persistent: completion is shared by deleting the row in source control, not by storing private state in one browser.

2026-08-05 application-boundary slice: the supporting-document lookup/open/path/view use cases now run through dashboard/supporting_document_application.py, with strict Pydantic request/response DTOs, ISupportingDocumentService, and injected SupportingDocumentPorts wired at the server composition root. The HTTP-facing names remain thin compatibility shims for existing handlers/tests; the policy and annotation workflow no longer lives in those endpoint bodies. server.py working file: 11,128 → 11,049 lines (-79, -0.71%; 11,049 versus 10,717 historical baseline, +332). The observed HEAD→worktree delta went from +4 lines/+339 bytes to -75 lines/-2,751 bytes; the attributable slice is -79 lines after separating the pre-existing concurrent working-tree edits. Three-commit trend remains +63, +0, +43 lines (latest first), so this slice reverses the current worktree growth. Focused supporting-document tests: 30 passed. System Python lacks pytest/Pydantic; validation used /home/adamsl/rol_finances/.venv-pytest/bin/pytest with PYTHONPATH=.. Next deletion: remove the four server-local compatibility wrappers once route/test callers are injected against ISupportingDocumentService. 2026-08-06 Walgreens annotation slice: the generic _expand_selected_row_region boundary is now applied by both the PDF and image annotators, so expense 2003 (WALGREENS #15466, 2025-01-17, $141.76) boxes the complete visual row (date, merchant/description, and amount) without absorbing the neighboring $9.85 Walgreens row. ANNOTATION_SCHEMA_VERSION advanced from 6 to 7 to invalidate narrow cached annotations. Regression coverage includes the image strategy and remains fail-closed for ambiguous row identity. server.py was not edited: 11,092 → 11,092 lines (0, 0.00%); measured SHA-256 59a525806b2b3b9e7484b8890369eecef7d2607525d984ace4f5c20619c58ef6. The last-three-commit trend remains +63, +0, +43 lines; HEAD→worktree is -32 lines/-524 bytes (11,124 → 11,092) from concurrent worktree edits, and the attributable slice is 0. Focused annotation tests: 7 passed; supporting-document tests: 5 passed. The broader annotation suite still has four environment-blocked image OCR failures because system tesseract is unavailable. The next deletion target remains the four server-local compatibility wrappers.
Status Interface / Port Object / Adapter Current seam Exit test
0 — Test and composition foundation
ITestEnvironmentDashboardTestEnvironmentVenv capability setup and test isolationFull suite is green or optional capabilities skip explicitly
Pydantic boundary contractStrictModelUntyped dicts, tuples, and optional fields crossing seamsStrict/frozen/extra-forbid models cover every new boundary
IPayloadCodecPydanticPayloadCodecManual JSON decoding, coercion, and response shapingNegative contract tests reject wrong and extra fields
ITypeCheckGatePyrightAndTypeScriptGateDashboard Python/JS currently outside strict project type gatesStrict Pyright and tsc --noEmit pass in one command
IContractTestSuiteSharedContractTestSuiteAdapter and fake behavior can drift independentlyEvery port implementation runs the same contract suite
IClockSystemClock, FixedClocktime.time(), datetime.now(), TTLsTime-sensitive services run against FixedClock
ICommandRunnerSubprocessCommandRunnerDirect subprocess.run/PopenServices assert commands without spawning processes
IHttpTransportUrllibHttpTransportDirect urllib.request callsTimeout, HTTP error, and JSON contracts have adapter tests
IActivityRecorderJsonActivityRecorder, NullActivityRecorderJSON log helpers and cross-cutting eventsNo service imports a concrete log writer
1 — HTTP boundary
IRouteControllerRouteRegistrydo_GET/do_POST condition chainsEvery API route is registered and duplicate routes fail startup
IRequestBodyReaderHttpRequestBodyReaderBody length, decoding, JSON parsingMalformed/empty/oversized bodies have contract tests
IResponseWriterDashboardResponseWriterjson_response, error_response, cache headersStatus, headers, and response shapes match characterization tests
IStaticAssetStoreLocalStaticAssetStoreHERE/REPO_ROOT file resolution and servingTraversal, type, missing-file, and no-cache tests pass
ITerminalSessionPtyTerminalSessionPTY spawn, resize, reap, WebSocket bridgeFrame/session lifecycle tested without a real shell
IDashboardApplicationDashboardApplicationStartup threads, caches, handler globalsserver.py only constructs and starts the application
2 — Letta agents
ILettaGatewayUrllibLettaGatewayAgents, messages, thoughts, tool calls; model reads migrated behind AgentModelOptionsService on 2026-07-30No agent service constructs a Letta URL; remaining callers keep this row active
IAgentDirectoryCachedAgentDirectoryLETTA_AGENTS, discovery, cache refreshRegistry ordering and cache behavior are unit tested
IAgentMessageServiceAgentMessageService/api/test and message shapingReset/send/fallback behavior uses a fake gateway
IAgentHealthProbeAgentHealthProbeRequired tools, SDK endpoint, send errorsHealth matrix is deterministic without live agents
IAgentPromptRunnerLettaCodePromptRunnerHeadless CLI command and prompt validationPermission mode, timeout, and final-result tests stay pinned. 2026-08-02 timeout-contract slice: a real Mazda turn outlived the former 330-second server budget after an internal SDK call consumed 180 seconds; the server and per-call browser budgets are now 900/930 seconds, and the prompt-runner test pins the server contract. The extraction remains open. server.py: 11,018 → 11,018 lines (+0, 0.00%; 11,018 versus the 10,486 historical baseline, +532). Three-commit trend: +220, +53, +0; HEAD→worktree remained +28 lines/+1,664 bytes because this slice changed a value without adding composition-root code. Next deletion target: move command construction, validation, subprocess execution, and result decoding behind the typed IAgentPromptRunner port.
IAgentSystemMessageRepositoryFileAgentSystemMessageRepositoryPer-agent system message filesMissing and unsafe paths fail closed
3 — Server Management
IHealthProbeHttpHealthProbe, TcpHealthProbe, CommandHealthProbeHEALTH_CHECKS and check variantsAdding a probe requires an adapter plus one registry entry
IServerRegistryServerRegistrySERVERS config and validationInvalid dependencies/check names fail at construction
IServerHealthServiceServerHealthServiceStatus computation, dependency roll-up, cacheAll green/yellow/red transitions have clock-controlled tests
IServerControllerServerControllerRestart/start command dispatchUnknown, blocked, starting, and success outcomes are tested. 2026-08-29 Agent Blocks startup slice: the embedded documentation SPA is now a declared dashboard startup dependency instead of a manually started hidden prerequisite. AgentBlocksServer owns health-before-launch idempotence, detached process construction, logging, and a strict start-result boundary; server.py only constructs it and exposes the compatibility command. RED was ModuleNotFoundError: servers.agent_blocks; focused verification is 147 passing tests, and the new module has 100% statement/branch coverage. server.py: 7,243 → 7,235 lines (-8, -0.11%; 7,235 versus the 10,486 historical baseline, -3,251, -31.00%). The observed worktree delta also changed by exactly -8 lines, so no concurrent movement affected this slice. 2026-09-06 ChatGPT Browser Server lifecycle slice: replaced the unmanaged remote nohup launcher with IBrowserServerLifecycle and SshBrowserServerLifecycle, a strict result model, injected health/command/clock collaborators, and an enabled browser-server.service on the Win10 WSL node. The dashboard Restart command now reaches the managed unit and verifies health instead of creating an orphan process. RED was ModuleNotFoundError: servers.browser_server; focused verification is 8 passing tests. server.py: 7,549 → 7,528 lines (-21, -0.28%; 7,528 versus the 10,486 historical baseline, -2,958, -28.21%). No concurrent server.py movement occurred, so observed and attributable deltas are both -21. This row remains open; next deletion target is moving the executor and Frita launch handlers behind the same typed server-controller boundary.
IServerLogSourceFileServerLogSource, RemoteServerLogSourceTailing, filters, remote Letta log pullSequence/filter behavior runs against fixtures
ISshGatewayOpenSshGatewaySSH health, Windows/WSL commandsNo health/controller object assembles raw SSH commands. 2026-08-06 address correction: the Windows 11 host moved to 100.118.122.75; the SSH registry and gateway contract fixtures now use that current Tailscale address. server.py: 11,063 → 11,063 lines (0, 0%); current working file remains 577 lines above the 10,486 historical baseline, with the last-three-commit trend still net-positive (+63, +0, +43) and the current worktree delta -61 lines versus HEAD. Next deletion target remains moving SSH target resolution behind ISshGateway.
IAvailabilityTrackerAvailabilityTrackerStarting windows and down duration globalsState transitions use injected clock and isolated state
IPcMetricsSourceLocalPcMetricsSource, SshPcMetricsSourcePC monitor extraction, cache, rate calculationParsing is pure and collection uses fake sources
4 — Model usage and provider health
IModelUsageSourceCodexUsageSource, ClaudeUsageSource, AntigravityUsageSourceToken extraction and remote source registryEach source passes one shared contract suite
IModelUsageServiceModelUsageServicemodel_stats, labels, classificationService contains no provider-specific condition chain
IUsageHistoryStoreJsonUsageHistoryStoreSamples, pruning, burn rate, slow leak historyCorrupt/missing history degrades safely
IProviderHealthMonitorChatGptProviderHealthMonitorZero-token polling and fleet error flagsOne probe updates the complete provider fleet deterministically. 2026-08-03 categorizer-health slice: mazda_categorizer_fallback_health() now filters the shared provider-health JSON to the categorizer chain, so chatgpt-oauth-vision:* failures cannot turn the LLM Provider Fallbacks tab yellow/red. The frontend parent-tab reducer now leaves unrelated per-server concern states on their own tabs instead of rolling them into Server Management; regression tests cover vision-only data, mixed provider data, and parent-tab isolation. server.py: 11,018 → 11,034 lines (+16, +0.15%; 11,034 versus the 10,486 historical baseline, +548). The positive delta is a compatibility guard while the shared event file remains multi-domain; next deletion target: move provider-health aggregation behind the typed IProviderHealthMonitor port and delete the server-side filter shim. 2026-08-24 SDK-token assignment slice: the read-only /claude_sdk_status credential probe now feeds a non-agent run_claude_code_sdk (Mazda) row in Model Stats → Agent Assignments; expired/missing/unreachable executor credentials render red and valid credentials show their UTC expiry. server.py: 8,592 → 8,602 lines (+10, +0.12%; 8,602 versus the 10,486 historical baseline, -1,884, -17.96%). No concurrent server.py movement was observed; all +10 lines are attributable to composition-root wiring. Focused tests: 48 Python and 2 JavaScript passed. 2026-08-24 SDK-account switch slice: added a persistent eg1972/rbarnesrol selector with atomic credential synchronization to Frita, and the assignment row now exposes the current account. server.py: 8,602 → 8,611 lines (+9, +0.10%; 8,611 versus the 10,486 historical baseline, -1,875, -17.88%). No concurrent server.py movement was observed; all +9 lines are attributable to account-selection wiring. Focused tests: 128 Python and 2 JavaScript passed. Next deletion target: move the assignment payload wiring behind a typed provider-health/assignment port and remove the compatibility re-export from server.py. 2026-09-06 SDK live-activity and quota-bar correction: the shared executor now exposes one bounded activity stream for every run_claude_code_sdk caller and the Server Management home renders separate Windows 98 Thoughts, Tool Calls, and Messages dialogs. Cancellation terminates the complete SDK process group, and a single-flight guard prevents interrupted retries from stacking token-consuming sessions. The SDK assignment now carries a scalar weekly percentage plus the usage-report retry timestamp instead of placing the tuple in the percentage field; rendering and proxy behavior live in focused modules. server.py: 7,547 → 7,549 lines (+2, +0.03%); 7,549 versus the 10,486 historical baseline is -2,937 lines (-28.01%). No concurrent server.py movement occurred, so observed and attributable deltas are both +2. The non-negative delta is limited to destructuring the existing usage result at the composition boundary and passing its two fields to the focused assignment builder. Next deletion target remains moving the assignment payload wiring behind a typed provider-health/assignment port and removing this server-local composition.
IProviderHealthMonitorChatGptProviderHealthMonitorLetta/W11 token validity and operator synchronization2026-08-29 provider-token visibility slice: the typed provider-account status contract now distinguishes valid, expired, stale/rejected, and unavailable Letta token copies from the current W11 Codex credential without requiring token-hash equality. Agent Assignments renders an explicit red failure instead of ?; Server Management again carries a dedicated Mazda Letta Provider Token tile; and the globally mounted account controller offers a Windows 98 Yes/No synchronization dialog once per detected bad-token incident. A successful W11 synchronization invalidates both dashboard caches immediately. server.py: 7,223 → 7,243 lines (+20, +0.28%; 7,243 versus the 10,486 historical baseline, -3,243, -30.93%). No concurrent server.py movement was observed during the measured slice; all +20 lines are attributable to composition-root credential loading, cache invalidation, and assignment-row wiring, while state policy, rendering, and account installation remain in focused modules. Focused verification: 126 Python tests and 15 JavaScript tests passed; full JavaScript suite 2,433 passed / 2 skipped; full Python suite 3,196 passed / 2 skipped with one unrelated pre-existing category-name mismatch. Next deletion target: move the assignment payload wiring behind a typed provider-health/assignment port and remove the compatibility re-export from server.py.
IProviderFailoverStrategyChatGptTokenFailoverStrategyStandby headroom, swap command, cooldownDecision logic and command adapter are independently tested
5 — Scanner and intake workflow
IScannerDeviceWindowsWiaScannerDevice selection, locking, WIA invocationScanner choice never depends on enumeration order
IScannerRegistryScannerRegistryWindow/Freezer configurationUnknown and duplicate devices fail at construction
IScannerDiagnosticsHpScannerDiagnostics2,232-line scanner/fix/diagnostic regionEvery LED and remediation message has fixture tests
IPrinterRepairServiceDeskJetPrinterRepairServicefix_deskjet_printerRepair never claims success without device evidence
IDocumentClassifierIntakeFacadeClassifierFacade invocation and classification shapingReady/busy/offline/error cases use one typed result
IDocumentStagerLocalAndRemoteDocumentStagerLocal copy and Win10 mirrorLocal authority and nonfatal mirror failure are pinned
IAgentConversationFactoryMazdaConversationFactoryPer-scan Letta conversation creationConversation failure stops dispatch without side effects
IDocumentDispatcherMazdaDocumentDispatcher482-line scan message and notify callMessage builder is pure; gateway owns delivery
ITrainerNotifierDetachedTrainerNotifier, NullTrainerNotifierTrainer command and detached spawnTests cannot spawn a real Trainer by construction. 2026-08-13 problem-only escalation slice: healthy scanner/PDF intakes now arm only an in-process deadline and do not launch a model session. Strict Pydantic TrainerLaunchRequest, IntakeCallback, TrainerEscalationResult, and TrainerEscalationNotice contracts cross the new ITrainerNotifier, ITrainerEscalationService, deadline, and escalation-recorder ports. Zero/incomplete/failed callbacks summon DetachedTrainerNotifier immediately; missing callbacks summon it after 900 seconds; valid stored or exact-duplicate callbacks cancel the deadline. Every launch attempt is persisted, pending watches are rebuilt from persisted processing intakes after a dashboard restart, and trainer_dispatched prevents duplicate escalation. The live systemd kill switch was re-enabled because the costly always-on policy is gone. server.py: 11,371 → 11,368 lines (-3, -0.03%); 11,368 versus the 10,486 historical baseline is +882 lines (+8.41%). No concurrent server.py movement occurred during the measured slice, so observed and attributable deltas are both -3. Focused suite: 19 passed. Next deletion target: extract the remaining process_scanned_document/process_pdf_document orchestration behind IDocumentIntakeService, removing the server-local watch/observe compatibility functions.
IIntakeEventStoreJsonIntakeEventStoreRecent report pointer and expense-stored eventsMerge, dedupe, pruning, and corrupt JSON are tested. 2026-07-30 slice completed: recent-intake event state now preserves statement archive_paths/archive_years so the UI can render the permanent filed scan path without re-deriving it first. 2026-07-30 slice completed: check-evidence siblings on the Last Scan / Recent Intake view now collapse into one visible expense row while preserving merged supporting-document slots. 2026-07-31 slice completed: the recent-pointer persistence now crosses an injected IIntakeEventStore boundary backed by an atomic JsonIntakeEventStore; a deterministic opaque snapshot token drives live Recent Report/Last Scan refreshes without coupling the browser to JSON files, and empty scanner outputs are persisted as visible failed intake attempts rather than leaving stale scanner tabs. 2026-09-06 recovered intake-state port slice: Windows 11's pre-existing WIP populated the typed IntakePort with intake_state_token(), allowing the Recent Report poll route to observe pointer-file changes without adding a new service-locator dependency. The accompanying Mazda tool reconciliation lives in focused Python and JavaScript modules rather than the composition root. server.py: 7,536 → 7,547 lines (+11, +0.15%); 7,547 versus the 10,486 historical baseline is -2,939 lines (-28.03%). No concurrent server.py movement occurred while the WIP was preserved; all +11 lines predated the synchronization and are attributable to the thin pointer-file token adapter. Focused verification: 191 Python and 13 JavaScript tests passed. Next deletion target: move the token implementation behind IIntakeEventStore and remove the server-local compatibility function. Next slice: move the remaining recent-intake/scanner-report row shaping and evidence-summary assembly out of server.py and into focused helpers/services so build_recent_intake_html() becomes mostly composition and presentation.
IDocumentIntakeServiceDocumentIntakeServiceprocess_scanned_document, process_pdf_documentWorkflow is an injected coordinator with no module globals
6 — ROL Finance
IReceiptReadStrategyReceiptReadServiceManual Circled Only, Total Only, and Several Expenses reads2026-08-30 manual-read Strategy slice: the former one-size-fits-all Mazda Fill path is gone. A strict ReceiptReadIntent selects one of three injected IReceiptReadStrategy objects: bounded circled-item and three-field readers, or the existing forensic document reader. Provider transport satisfies IFocusedReceiptReader; prompts, runtime schemas, provider adapters, the subprocess boundary, browser Commands, and progress controls each live in focused files under 161 lines. server.py: 7,235 → 7,242 lines (+7, +0.10%); 7,242 versus the 10,486 historical baseline is -3,244 lines (-30.94%). No concurrent server.py movement was observed; all +7 lines are attributable to composition-root Strategy construction. The positive delta is required wiring while 298 lines of old service code left the former module. Live Gemini checks: Total Only returned merchant/date/amount in 2.52 seconds; Circled Only returned the boxed $7.18 item in 11.44 seconds; a cropped receipt with no printed total failed closed. Next deletion target: route this application service through a typed http_app port and remove the remaining server-local endpoint shim.
IExpenseReportSynchronizerStaticExpenseReportSynchronizerStored expense edits versus static report rows2026-08-27 static-report synchronization slice: Edit Expense now invokes an injected report adapter after the database write. It resolves legacy report-only rows by their previous date/amount/description identity, stamps the recovered expense ID, and rewrites only the selected static transaction row; the same vendor's other dates remain untouched. server.py: 7,178 → 7,191 lines (+13, +0.18%); 7,191 versus the 10,486 historical baseline is -3,295 lines (-31.42%). No concurrent server.py movement was observed; the non-negative delta is composition-root import/invocation and fail-loud warning wiring, while matching and filesystem behavior live in finance/expense_report_sync.py. Focused verification: 21 tests passed. Next deletion target remains routing expense commands through the typed ExpensePort and removing server-local orchestration.
IExpenseEditAuditLogJsonlExpenseEditAuditLogPer-request Edit Expense diagnostic evidence2026-08-30 Edit Expense audit slice: every edit attempt now crosses an injected audit-log port and an AuditedExpenseEditCommand Decorator. The persistent mode-0600 JSON Lines adapter records the allowlisted request, exact JSON success/failure result, changed fields, warnings, returned record, UTC time, and action ID; an audit-disk failure cannot change or conceal the database command's authoritative result. Pytest injects the null adapter, so test edits never contaminate the live audit. server.py: 7,242 → 7,263 lines (+21, +0.29%); 7,263 versus the 10,486 historical baseline is -3,223 lines (-30.74%). The pre-existing staged receipt-read refactor was already present in the 7,242 start count; no concurrent movement affected this measured slice, so the attributable and observed slice deltas are both +21. The positive delta is limited to composition-root imports, adapter construction, and wrapping the legacy command; persistence, filtering, failure isolation, and the Decorator live in finance/expense_edit_audit.py. Focused verification: 120 tests passed. Next deletion target: populate the typed ExpensePort, move the remaining edit orchestration out of server.py, and delete the compatibility command wrapper.
IReceiptDestinationPolicyCanonicalReceiptDestinationPolicyEdited receipt filename, day folder, database references, and scanner pointer2026-08-30 canonical receipt-refiling slice: corrected dates now move receipt images through an injected Strategy to the complete year/month/month_DD destination instead of renaming only inside the old day folder. An IExpenseReceiptSynchronizer Bridge updates id_light, receipt_url, matching document_url, source_file, and missing Recent Report archive_paths together; collisions remain fail-closed. The oversized repository was split from 381 to 249 lines, and the 749-line test file became focused files no larger than 151 lines. server.py: 7,263 → 7,270 lines (+7, +0.10%); 7,270 versus the 10,486 historical baseline is -3,216 lines (-30.67%). No concurrent movement affected this measured slice; observed and attributable deltas are both +7, limited to Strategy construction and injection. Verification: 3,192 Python and 2,441 JavaScript tests passed (2 intentional skips in each suite). Next deletion target: inject the policy through the typed ExpensePort and remove the server-global construction.
IExpenseRepositoryMySqlExpenseRepositoryExpense lookup, category changes, notesServices never import DB connection helpers. 2026-08-26 Verified Transactions actions slice: the Last Scan table now exposes Edit, confirmed Delete, and fixed Michigan 6% tax commands. Stored-row reads/deletes remain behind IExpenseRecordRepository/MySqlExpenseRecordRepository; exact Decimal tax arithmetic lives in finance/sales_tax.py; the browser synchronizes row removal and taxed amounts with the existing Prev/Next expense dialog through a declared mounted-widget registry. server.py: 7,456 → 7,469 lines (+13, +0.17%) during the completion slice; 7,469 versus the 10,486 historical baseline is -3,017 lines (-28.77%). The interrupted feature's complete HEAD→worktree movement is +91 lines (7,378 → 7,469); no concurrent server.py movement was observed during this completion slice. The non-negative slice delta is thin command orchestration and compatibility exports; database behavior, tax policy, and browser behavior live outside the composition root. Next deletion target: populate the typed ExpensePort in http_app/ports.py, inject these three expense commands into the POST registry, and remove their server-local compatibility functions. Focused verification: 259 Python and 58 JavaScript tests passed. 2026-08-26 receipt-image synchronization follow-up: Edit, Delete, and Add 6% now invoke an injected RecentReportImageSynchronizer after the expense write. It finds the matching shared/per-scanner intake, recomputes the total from all remaining document rows with Decimal, regenerates the established vendor/date/amount stem, renames the archived image without overwriting a collision, and persists the new archive path. server.py: 7,469 → 7,513 lines (+44, +0.59%); 7,513 versus the 10,486 historical baseline is -2,973 lines (-28.35%). No concurrent server.py movement was observed during this follow-up, so observed and attributable deltas are both +44. The non-negative result is composition-root invocation and fail-loud response wiring; naming, aggregation, filesystem behavior, and pointer mutation live in finance/recent_report_image.py. Focused verification: 562 Python and 58 JavaScript tests passed. Next deletion target remains injecting these commands through ExpensePort and deleting their server-local endpoint orchestration.
IExpenseRepositoryRecentReportImageSynchronizerExplicit vendor corrections versus same-receipt line-item additions2026-08-26 vendor-identity correction: explicit stored-expense edits may replace an incorrectly archived vendor identity, while same-receipt additions continue preserving it. server.py: 7,532 → 7,533 lines (+1, +0.01%); 7,533 versus the 10,486 historical baseline is -2,953 lines (-28.16%). No concurrent movement was observed; the attributable +1 is a composition-root option passed to the focused image synchronizer. Next deletion target remains routing expense commands through the typed ExpensePort.
IHumanVerificationRepositoryMySqlHumanVerificationRepositoryPersist and present human review of expense rows2026-09-10 human-verification slice: opening Set Category sends a strict expense-ID command through the populated CategoryPort to an injected repository. The additive expenses.human_verified migration defaults every row to false; the adapter changes it idempotently to true, static reports hydrate it in memory, and dynamic rows carry it directly. Picker-owned CSS renders a green lower-right triangle without rewriting report files. RED was ModuleNotFoundError: finance.human_verification plus three pipeline assertions. server.py: 7,568 → 7,595 lines (+27, +0.36%); 7,595 versus the 10,486 historical baseline is -2,891 lines (-27.57%). No concurrent movement occurred; observed and attributable deltas are both +27. The non-negative delta is composition-root construction, compatibility commands, and selecting the field in existing dynamic builders; validation, persistence, hydration, browser behavior, and styling live in focused modules or the report injector. Focused verification: 271 dashboard and 91 pipeline tests passed. The full dashboard run passed 3,352 tests with 2 skips and reproduced 3 unrelated baseline failures against untouched HEAD. Next deletion target: populate the remaining CategoryPort methods and remove server-local category/verification wrappers.
IReceiptRepositoryFileReceiptRepositoryReceipt mounts, index, matching, URL mappingMount/index policy has contract tests and injected clock
IReportRepositoryFileReportRepositoryReport aliases, discovery, status, row lookupPath safety and ambiguity rules are isolated
IReportRowWriterHtmlReportRowWriterStatic row category/color updatesWrites are idempotent and fixture snapshots stay stable
IExpenseCategorizationServiceExpenseCategorizationServicerecategorize_expense and undo orchestrationDB, report, taxonomy, undo ports are constructor-injected
IReceiptLookupServiceReceiptLookupServiceLookup/presence/source-document resolutionEvery matching tier is tested without real files or DB
ISupportingDocumentServiceSupportingDocumentServiceSupporting-document lookup/open flowUses existing annotation port and injected repositories. 2026-07-30 slice completed: source-document resolution no longer silently falls back to the receipt path when no distinct source document exists, so View Source Document and View Receipt cannot collapse onto the same file. 2026-07-30 slice completed: receipt-vs-source equivalence detection now lives in a focused supporting-document helper so scanner/recent-intake dialogs suppress the source button when document_url and receipt_url resolve to the same underlying file through different spellings. 2026-08-02 slice completed: the expense lookup boundary now adapts to the deployed receipt-only expenses schema instead of selecting optional dashboard columns unconditionally; absent id_light, document, statement, ledger, notes, and role fields are exposed as stable nulls, allowing the existing annotation port to reach an available receipt and draw its red box. Regression coverage pins the reported expense shape (1120 / $53.06) against this schema, plus existing highlighted viewer coverage. 2026-08-05 slice completed: scanner/recent-intake paper scans are now offered only through View Scanned Statement; View Source Document requires a distinct downloadable document_url or real report source and suppresses equivalent references. The source/scanned fallback policy remains in server.py pending extraction into a dedicated supporting-document service. Regression coverage includes the Last Freezer Scan path and same-file suppression. server.py: 11,061 → 11,101 lines (+40, +0.36%; 11,101 versus 10,717 historical baseline, +384). The working-file delta includes the schema-compatibility correction and this slice; the focused document-policy change is net-positive because it adds an explicit scanned-statement fallback and fail-closed source selection. Three-commit trend: +220, +53, +0; HEAD→worktree after this slice is +40 lines/+1,794 bytes. 2026-08-05 slice completed: that follow-up is done for the page-policy half. Report-page identity is now a strict discriminated union (finance/report_page.py: ScannerReportPage, RecentIntakePage, RecentReportPage, MonthReportPage) parsed once instead of re-parsing the report URL inside every resolver, and the fallback rules moved behind IIntakePageLookup + SupportingDocumentPageResolver (finance/supporting_documents.py), constructor-injected and tested with a fake port — no server, filesystem, or DB. The three call sites that each re-implemented “an empty scanned statement falls back to the page scan” now share slot_reference(), driven by a new falls_back_to_page_scan flag on the slot catalog, so a fourth slot is a data edit. The schema-compatibility shim is likewise gone: both _matching_expense and _fetch_expenses_by_ids now share finance/expense_schema.py (ExpenseSchema strict model, IExpenseSchemaProbe port, ShowColumnsProbe/InformationSchemaProbe adapters held to one parameterized contract suite), which turned test_fetch_expenses_by_ids_supports_minimal_live_expenses_schema green. StrictModel now has one home in dashboard/contracts.py. server.py: 11,101 → 11,124 lines (+23). The positive delta is the composition adapter (_ServerIntakePageLookup + the lazy resolver factory) that policy movement requires; 43 lines of probe/parse/fallback logic left the file and 220 lines of new domain code live outside it. Suite: 657 passed, 18 pre-existing failures unchanged (missing tesseract, plus tests naming inspect_scan_image_quality, which does not exist in this checkout). 2026-08-05 annotation/supporting-document regression slice: the Diners Club expense 2000 fixture now exercises OCR aliases and compact/leading-artifact amount forms while the production annotator keeps the row identity and amount checks fail-closed; the supporting-document regression uses a temporary scan fixture and proves the downloaded PDF remains View Source Document while the JPG remains View Scanned Statement. server.py was not edited in this slice: 11,124 → 11,124 lines (0, 0.00%; 11,124 versus 10,717 baseline, +407); the observed working-file delta and slice-attributable delta are both 0. Three-commit trend remains +63, +0, +43 lines (latest first); HEAD→worktree is 0 lines/0 bytes. Targeted supporting-document tests: 30 passed; freezer annotation regression: 1 passed. The broader annotation suite has 20 passed and 4 environment-blocked failures because this machine has the Python pytesseract wrapper but no system tesseract executable; the next cleanup remains moving _source_document_reference behind ISupportingDocumentService, which should remove the remaining server-local resolver shim rather than adding more compatibility code.
IRecentIntakeEventRouterExactRecentIntakeEventRouterSTEP 8 callback identity and Recent Report target selection2026-09-10 fail-closed callback-routing slice: uncorrelated event-bus callbacks can no longer fall through to whichever Window/Freezer intake happens to be latest. A strict Pydantic RecentIntakeEventIdentity owns document/conversation/dispatch evidence, the ABC port declares target selection, and the concrete Strategy preserves exact dispatch matching plus path tie-breaking. server.py: 7,612 → 7,568 lines (-44, -0.58%); 7,568 versus the 10,486 historical baseline is -2,918 lines (-27.83%). No concurrent server.py movement was observed, so the observed and attributable deltas are both -44. Focused routing/status verification: 14 tests passed. Next deletion target: move _fold_event_into_intake and duplicate-ID recovery behind a typed intake-event application service.
IFinanceReportServiceFinanceReportServiceRecent intake/report and receipt-only HTMLRendering is separate from data acquisition. 2026-07-30 slice completed: recent-intake HTML now shows a distinct Archived Scan Copy line and prefers stored archive evidence before falling back to derived lookup. 2026-07-30 slice completed: the synthetic Last Scan / Recent Intake table now renders one visible expense row for a real-world expense even when separate check evidence exists. 2026-08-20 report-source viewer slice: report pages can now open their exact PDF/image/workbook source through a dedicated read-only route; Excel sources reuse the existing browser renderer and generated report HTML is rejected as a source. server.py: 12,430 → 12,468 lines (+38, +0.31%); 12,468 versus the 10,486 historical baseline is +1,982 lines (+18.90%). No concurrent server.py movement occurred during the measured slice, so observed and attributable deltas are both +38. The positive delta is the thin composition-root route and compatibility helper needed to expose the already-existing source resolver; rendering remains delegated to the existing document adapter. Next deletion target: move the route response/content-type assembly behind IFinanceReportService and remove the server-local helper.
IRecentScanRepositoryMySqlRecentScanRepositoryRecent scans and month status queriesQuery results map to typed records behind one port
IVendorReviewServiceVendorReviewServiceVendor keys, pending review, set vendorValidation, categorization, and persistence are isolated

Established examples — do not move backward

These are outside the active count because they already demonstrate the intended shape.

ICategoryTaxonomyStatic, MySQL, and fallback implementations in category_taxonomy.py.
Category undo portsRepository and store contracts with compare-and-swap orchestration in category_undo.py.
IExpenseDocumentAnnotationServiceFormat strategies and composition factory in document_annotation.py.
Voice pipelineTranscription, cleanup, Letta adapter, pipeline, and edge-tts synthesis strategy separated under voice/.
Statement reviewReview behavior already isolated in statement_review.py.
Frontend GoF layerInterfaces under js/abstract/, concretes under js/implementation/.

One-slice workflow

  1. Choose exactly one active row; record its current callers and mutable globals.
  2. Write the failing consumer test first, run it, and save the RED command plus failure.
  3. Add characterization, shared-contract, and negative-validation tests before production code.
  4. Define the smallest fully typed ABC and strict Pydantic models in the owning domain. The abstract module imports no concrete adapter.
  5. Implement only enough typed behavior for GREEN, including a minimal fake that passes the same contract tests.
  6. Inject the port into the service/controller; construct the concrete only at the composition root.
  7. Leave a temporary forwarding function in server.py when compatibility requires it.
  8. Move relevant tests from test_server.py into the matching domain test file.
  9. Refactor with tests green, then run Pyright, TypeScript, focused tests, full suites, coverage, and the relevant manual smoke check.
  10. Record current line counts, delete this work row, and commit the shrinking plan with the implementation.

Required evidence in each handoff or PR

Slice:
Observed RED:
RED command and failure:
Interface:
Pydantic model(s):
TypeScript type(s):
Concrete Object(s):
Callers migrated:
Compatibility shim:
Focused tests:
Contract/negative tests:
Pyright result:
TypeScript result:
Branch coverage:
Full Python result:
Full JavaScript result:
server.py lines before/after:
test_server.py lines before/after:
Ledger row removed:

Decision log

2026-07-29
Use an incremental ports-and-adapters rewrite; do not replace the stdlib HTTP server or change routes as part of extraction.
2026-07-29
Interfaces are required at meaningful boundaries and all cross-cutting side effects. Pure transforms remain functions.
2026-07-29
Completed work rows are deleted instead of archived here, so this page becomes smaller as the rewrite advances.
2026-07-29
Begin with the green test baseline, then model usage as the pilot production slice.
2026-07-29
Adopt strict Pydantic v2 models for every Python boundary and strict TypeScript for all new frontend work; static and runtime validation are both required.
2026-07-29
Adopt strict TDD: witness RED before production code, reach GREEN minimally, then refactor; new domain/application modules require 100% branch coverage.
2026-07-30
Completed the model-read sub-slice of ILettaGateway: RED was pytest -q tests/test_letta_gateway.py failing because the agents boundary did not exist; GREEN was 14 focused gateway/model-option tests. Added strict frozen Pydantic contracts, FakeLettaGateway, UrllibLettaGateway, composition-root wiring, and a temporary agent_model_payload forwarding shim. The ledger row remains because messages, thoughts, tool calls, and write paths are not migrated.
2026-07-31
Made server.py size a mandatory rewrite metric. Historical baseline: 10,486 lines. The pre-model-slice working-file backup was 10,871 lines; counting only the model-read extraction hunks gives 10,871 → 10,877 (+6, +0.06%) because composition-root wiring and compatibility shims were added before all legacy paths could be deleted. The current working file is 10,962 lines (+476, +4.54% versus baseline); unrelated concurrent dashboard work accounts for the remaining +85 lines in the observed +91 since the backup. Future slices must capture immediate before/after counts and report observed and slice-attributable deltas separately.
2026-08-26
Fixed the same-receipt rename boundary: RecentReportImageSynchronizer now preserves the vendor/date identity already stamped on the archived receipt, changes only the aggregate amount, and updates every associated expense's receipt_url/source_file. Last Scanner pages also reconstruct Archive Verification on mount instead of showing it only momentarily after Save All. The selected slice added 31 composition-root lines; concurrent rewrite work reduced other imports/wrappers while this fix was in progress, so the observed working file is 7,532 lines. Current count versus the 10,486 historical baseline is -2,954 lines (-28.17%). The non-negative attributable delta is the injected database-reference adapter; filename identity, aggregation, and rename behavior remain in finance/recent_report_image.py. Next deletion target: move that adapter behind the existing expense port and remove the server-local compatibility function.