Living Plan · Round 13 landed 2026-08-26 · Round 14 brief

Cutting server.py to the Bone

The target has moved. dashboard/server.py is no longer being halved — it is being reduced to a composition root of under 2,000 lines, and designed to land near 700. The method changes with it: stop carving off whichever cluster measures best this week, and instead dismantle the three structures that are holding 7,500 lines up. Many small Pydantic models. Many small JavaScript interfaces. Patterns with names.

Two of the three are down. Round 12 put ports between the routes and server, so moving code costs a port method instead of a permanent re-export. Round 13 spent that: every hand-maintained config literal is gone, module-level assignment fell from 704 lines to 262, and the roster it typed turned out to be serving a duplicate agent tile to the live dashboard. What is left of the three is the six hand-rolled caches — round 14, and the thing the whole agent block is pinned behind.

11,407 → 7,147 lines new target: < 2,000 3,162 tests passing last landed: round 13 = d04b519d srv.: 220 → 134 sites config literal: 430 → 0 lines

Start here

You are picking up round 14 — one cache class, six instances. Do these six things before writing any code.

  1. Run the sync-all skill. Three boxes, one branch, and the dashboard you are changing is served from a machine you are probably not typing on. Then run git status again when you are about to commit. Round 12 was mid-flight when a concurrent agent ran git add -A and swept two of its unfinished files into an unrelated commit. Nothing was lost, but the tree is genuinely shared and “I checked at the start” is not the same as “it is still true.”
  2. Read Round 14 in full. It removes the least of any round on the board (~180 lines) and unblocks the most: the agent cluster is six caches deep and round 24 cannot start until they are one class.
  3. Read What round 13 actually did. You are inheriting its registries, two of its findings, and one habit worth copying: it read the old literals out of git and diffed them against the new derived views on live values, so the golden test could not decay into a restatement of the new code.
  4. Read What round 12 actually did. You will be adding to the mechanism it built, and it left one trap that is invisible until it bites: DIRECT_SERVICES in tests/http_app_harness.py. Round 13 needed no entry there, because it routed everything through ports rather than at owning modules — but it hit the same trap twice from the other side, and its notes say where.
  5. Read The three structures. The first two are done; the third is your round.
  6. Re-run the census (the commands are here) against whatever server.py actually is when you arrive. Every number on this page is from d04b519d and every number on this page will be stale for you — a second agent ships features into this file while you work, and round 12 measured a 29-line growth across a round that removed 41.
  7. Read the standing rules. Rules 1–11 were paid for in production incidents; 12–17 came with the retarget and two of them amend earlier ones; 18–19 came from round 12. Rule 10 in particular has been deliberately relaxed.

Why the target moved from 5,700 to 2,000

The halfway target was reachable and the plan was on course for it. That is exactly the problem: 5,700 lines of procedural code with no classes in it is not a better building than 11,407 lines of the same. It is the same building with a third of the rooms in the garage. The rounds were getting slower and smaller — round 11 measured 153 lines and removed 107 — because the cheap, cohesive clusters are gone and what remains is held in place by structure, not by size.

Round 11's own recommendation was to stop taking small slices and spend a round designing a seam. That instinct was right and the scope was too modest. The retarget makes it explicit:

The old question

“Which cluster can I move this week with the fewest names left behind?”

Answer quality degrades every round by construction: you are always picking from a set that gets worse.

The new question

“What is server.py for, and what has to be true before everything that isn't that can leave?”

Answer: it is a composition root. About 700 lines. Everything else is someone's module.

Under 2,000 is the ceiling, not the design. The design is itemised below and comes to roughly 700. The 1,300 lines of daylight between them is the slack this plan has learned, over eleven rounds, that it always needs.

The census — what 7,147 lines actually are

Measured on d04b519d. Reproduce it from dashboard/ with the AST script in How to rank a cluster plus these four one-liners:

grep -oh 'srv\.[A-Za-z_][A-Za-z0-9_]*' http_app/*.py | sort -u | wc -l   # 111 (109 + 2 docstring artefacts)
grep -c 'return {' server.py                                            # 162
grep -o '\.get(' server.py | wc -l                                      # 473
grep -c 'with .*_lock' server.py                                        # 36

Report your round by its diff, not by wc -l. Round 13 landed −579/+192 in server.py and the file moved 7,533 → 7,146 — those do not reconcile, and they are not meant to. Two agents write to this tree: round 12 removed 41 lines while a parallel feature added ~170 on the same days, and round 13 arrived to find the srv. ceiling test already red for the same reason.

0

classes defined in 7,147 lines. Not “few”. Zero. Every piece of state in this file is a module global and every piece of behaviour is a top-level def. Rounds 12 and 13 added classes to the HTTP side and to the new registry modules; this file still has none of its own.

241

top-level functions, 5,396 lines of bodies. 25 of them are over 50 lines and account for 2,093 — 28% of the file lives in 10% of the functions. Round 13 moved data, so this number did not move at all. Rounds 15–21 are what shift it.

109

distinct names http_app/ still reaches into server for, across 134 call sites — down from 167 / 220 before round 12 and 117 / 148 before round 13. 22 of what is left are private. What round 13 did.

262

lines of module-level assignment, down from 704. The ~430 lines of config literal are gone — SERVERS, AGENT_CARDS, LETTA_AGENTS, SCANNERS, the four REPORTING_CATEGORY_* dicts and six smaller ones are now typed collections in seven modules. Round 13.

20

module-level threading.Lock() objects, 36 with …_lock: blocks, 7 hand-rolled TTL constants, and six independent hand-written read-through caches — while monitoring/win10_node.py already contains a tested _TtlCache class.

162

return { statements — untyped payload boundaries, most of them going straight to the browser — against 473 .get() calls, each one a place a missing key becomes a default instead of an error. The dashboard now has 48 StrictModel subclasses — round 13 added 13.

The three structures holding the file up

Eleven rounds of extraction did not shift any of these, because none of them is a cluster. Each one has a name in the literature and a standard remedy. Two are now done; the third is round 14.

1. srv is a Service Locator anti-pattern mechanism built, round 12

http_app/services.py exposes srv, a PEP 562 module __getattr__ onto server. The two route ladders reached through it for 167 distinct names across 220 call sites, 26 of them private (srv._fetch_month_status, srv._load_json, srv._voice_log_lock — a lock, reached across a module boundary, by an HTTP route).

This is why every round paid a re-export tax. Move a function out and the route still says srv.the_name, so the name has to stay behind as an import. 47 names in server.py existed for no other reason, and they would never have expired on their own, because their caller was the one thing not moving.

The remedy: Facade + Registry. Round 12 built it — http_app/ports.py and http_app/registry.py — converted all 47 free names, and converted ScannerPort end to end as the worked example. Round 13 took the ceiling to 109 names / 134 sites, and tests/test_http_app_ports.py pins it as a ceiling that only ever falls.

What is left is the interesting half. The remaining 109 are real services, not re-exports, and they leave one at a time as their round populates its port. Your job includes lowering that ceiling by the names your block owns — and the mechanism is working: round 13 caught a parallel feature that had added a name without lowering anything, because the ceiling test was red before a line of round-13 code was written.

2. Configuration is source code, and nothing validates it done, round 13

About 430 lines were hand-maintained data literals: not logic, never typed, and several of them parallel — multiple dicts keyed alike that had to agree, with nothing checking that they did. They are now seven typed registry modules, and round 13 is the record of what that found. The diagnosis below is kept as written, because two of its predictions are worth reading against what actually happened: the REPORTING_CATEGORY_* family turned out to be a defence exactly as predicted, and SERVERS turned out to be clean — while LETTA_AGENTS, which this section never suspected, was serving a duplicate agent tile to the live dashboard.

Verified, and it was the round-13 headline. REPORTING_CATEGORY_DB_MAP, REPORTING_CATEGORY_CLASS and REPORTING_CATEGORY_STYLE are three dicts keyed by the same thirteen category names, and REPORTING_CATEGORY_ANCESTOR_MAP is a fourth keyed by the ids the first one hands out. Adding a category to one and forgetting another produces a report row with no CSS class or no colour — rendered, served, and indistinguishable from a styling choice. One ReportingCategory model with a name / db_id / css_class / background / font shape, one list of thirteen, and the four dicts become derived views that cannot disagree. 58 lines become about 20, and a whole class of silent defect stops being expressible. Read the comment above REPORTING_CATEGORY_ANCESTOR_MAP before you start: these four are now the offline seed for LEGACY_TAXONOMY, superseded at runtime by ICategoryTaxonomy. That makes them safer to restructure, and it also means the validator you add is guarding a fallback path — say so, per rule 11.

SERVERS (159 lines) and AGENT_CARDS (136) are the same shape of risk at four times the size: a mistyped systemd unit name in SERVERS is a Restart button that reports success and does nothing — the exact failure signature the round-6 postscript exists to warn about.

3. The same cache is written six times, by hand Template Method this is round 14

_agent_list, _agent_activity, _agent_health, _letta_id, _weekly_remaining and _model_stats_agents each carry a {'value': None, 'ts': 0.0} dict, a threading.Lock(), and a *_CACHE_TTL constant, with the read-through logic re-typed inline at every call site. build_agent_list adds a stale-while-revalidate background refresh to its copy and stores the in-flight flag in the same untyped dict.

monitoring/win10_node.py already has the class this wants — _TtlCache, with Win10CacheEntry as a typed entry and a deliberate design note about why a cached None must still count as a hit. It is tested. It has never been reused. Round 14 promotes it to caching.py, adds a StaleWhileRevalidateCache subclass for the one call site that needs it, and deletes six hand-rolled copies.

The destination: what server.py is allowed to be

Design the end state first, then subtract toward it. When the work is done, server.py contains this and nothing else:

PartLinesWhat it holds
Module docstring + imports~120 One import per owning module. No from x import name re-exports — those are what round 12 abolished. 38 gone, 129 srv names to go.
The container~200 Construct ~22 services with their real collaborators. One def build_*_service() each, resolved per call, never captured.
The port registry~80 Bind the dozen facade objects the route ladders depend on.
Startup + background loops~80 startup_tasks, startup_banners, and the thread starts that take a sweep callable so the bundle is rebuilt per iteration.
HTTP bootstrap~40 Port, handler class, serve. Everything else is in http_app/ already.
Blank lines, section comments, the design notes worth keeping ~180 This file has good comments. Keep the ones that explain a decision; drop the ones that narrate code that has left.
Total~700 Against a 2,000-line ceiling. The gap is the slack.

Anything you find yourself wanting to keep in server.py that is not on that list belongs somewhere else. That is the entire test.

The pattern kit

Seven patterns cover everything below. Name the one you are applying in the commit message — it is how the next shift knows whether you were moving code or changing shape.

PatternWhere it goesWhat it kills
Facade A dozen ports between http_app/ and the modules. 169 free names reached through srv, and the re-export tax on every future round.
Registry / Abstract Factory SERVERS, SCANNERS, RESTART_HANDLERS, HEALTH_CHECKS, AGENT_CARDS become typed collections built by a factory. ~430 lines of literal, and the whole family of “entry N is missing key K” bugs.
Template Method caching.py: TtlCache and StaleWhileRevalidateCache. Six hand-written caches, 14 of the 20 module locks, 6 TTL constants.
Chain of Responsibility Document resolution — _source_document_reference, _source_document_path, _resolve_local_supporting_document, _usable_document_reference, _slot_reference and friends are already a try-this-then-that chain written as nested if. ~610 lines and about 35 interlocking private helpers, replaced by an ordered list of small resolvers.
Command RESTART_HANDLERS / RESTARTABLE_KEYS, and the GET/POST route ladders themselves. A 23-line dispatch dict and two ~600-line if ladders.
Strategy Statement preflight and document processing — run_statement_preflight (157 lines) and process_scanned_document (171) are both one-function-per-document-kind wearing a switch. The two largest functions in the file.
Repository / Adapter The intake record's persistence layer, and every direct _rol_get_connection() reach from a render path. The reason notes has six stays for three functions: rendering code talking straight to a DB.

The burn-down — rounds 12 to 25

Fourteen rounds. The order is not by size; it is by what unblocks what. Rounds 12–14 remove almost nothing and make every later round cheaper. Take them in order.

The Left column is arithmetic on the round-11 baseline of 7,378. It is not a prediction of wc -l, because a second agent ships features into this file in parallel — round 12 removed 41 lines and the file still ended 154 higher; round 13 removed 579 and the file fell 387. Track the column as “refactor debt retired”, and report your round's diff separately from the file size.

#RoundPatternOut LeftWhy here
12Ports & the death of srv landed 6c69a3f4 Facade417,337 Enabling, and it worked: moving code now costs a port method, not a permanent re-export. srv. 220→147 sites. What it did.
13Config out, typed landed d04b519dRegistry 5796,758 Came in at 579 removed against ~530 predicted — the first round to beat its estimate, because round 12's ports meant the wiring went into a port method instead of back into the namespace. Found a live duplicate agent tile. What it did.
14One cache class, six instancesTemplate Method ~1806,580 Removes the least on the board and unblocks the most: the agent cluster is six caches deep and round 24 cannot start until they are one class. Round 14 in full.
15Intake record persistenceRepository ~4806,100 Round 11 already mapped this surface. _fold_event_into_intake (122) is the largest single piece.
16Intake & report HTMLStrategy ~5405,560 Needs 15 first. build_recent_intake_html is 171 lines of string building in a server module.
17Statement preflightStrategy ~4505,110 run_statement_preflight is 157 lines; the doc-kind switch is the seam.
18Document processingStrategy ~3904,720 process_scanned_document 171 + process_pdf_document 57 + reprocess_report 37, one strategy per kind.
19Document resolution chainChain of Responsibility ~5804,140 The densest knot in the file: ~35 private helpers, none over 70 lines, all calling each other.
20Receipt lookup & report rowsRepository ~5703,570 Absorbs the deferred open-receipt and notes clusters, which were never clusters of their own.
21Expense entry, edit, resolveRepository ~6202,950 _fetch_expenses_by_ids is 118 lines and submit_manual_receipt_entry 81.
22Scanner hardware, trainer, dispatchRegistry + Facade ~6002,350 Round 11 drew half this boundary already. Finishes the Mazda block.
23Server management & restart registryCommand ~7001,650 The ceiling falls here. SERVERS left in round 13; this is the behaviour that read it.
24Agent registry, health, model statsFacade ~6501,000 Only tractable after 14. Six caches, two locks, three payload builders.
25Report status & the sweepStrategy ~300~700 _extract_report_attention_detail (86) and whatever is still standing. Then delete the dead comments.

Read those numbers as content, not as yield. Historically a move returns 70–100% of its measured figure because wiring lands back in server.py. That ratio is exactly what round 12 attacks: once the routes talk to ports, the add-back goes into the container as one registration line instead of into the module namespace as a permanent name. If the burn-down is running behind by round 16, the diagnosis is almost certainly that round 12 was done partially. It was not: the mechanism is complete and the ceiling test will tell you the moment it starts leaking.

Round 14 in full — one cache class, six instances Template Method

About 180 lines, and the smallest removal on the board. Take it anyway, and take it next: round 24 (the agent registry, health and model stats) is the second-biggest block left and it is pinned behind this one. Six hand-written read-through caches sit under it, each with its own dict, its own lock and its own TTL constant, and no round can restructure that cluster while the caching is re-typed inline at every call site.

What is actually there

Six copies of the same eleven lines. _agent_list, _agent_activity, _agent_health, _letta_id, _weekly_remaining and _model_stats_agents each carry a {'value': None, 'ts': 0.0} dict, a threading.Lock() and a *_CACHE_TTL constant. That is 14 of server.py's 20 module-level locks and 6 of its 7 hand-rolled TTL constants. Round 13 did not touch any of them — it moved data, and these are state.

build_agent_list is the odd one and the interesting one. It adds stale-while-revalidate on top: when the entry is stale it serves the stale value immediately and kicks a background thread, because a cold rebuild can block >10s on the Letta roster fetch and trip the browser's fetch timeout. It stores the in-flight flag in the same untyped dict as the value, so a rebuild that raises leaves refreshing stuck True and the cache never refreshes again.

The class already exists, tested, and has never been reused

monitoring/win10_node.py contains _TtlCache with Win10CacheEntry as a typed entry, and a deliberate design note about why a cached None must still count as a hit — the entry, not the value, is the presence sentinel. Five of the six hand-rolled copies get that wrong: they test value is not None, so any cache whose legal answers include None re-fetches every single call while looking exactly like a working cache.

Take them in this order

  1. Promote _TtlCache to caching.py unchanged, with CacheEntry as the typed entry, and leave monitoring/win10_node.py importing it. Do this as its own step and keep tests/test_win10_node.py green before you convert anything else — it is the only proof you have that the class behaves, and it is proof written against the real caller.
  2. Convert the five plain caches. Mechanical. _letta_id is the one to do first: it is the one whose None is a legitimate cached answer (an agent absent from the Letta roster) and it already carries LETTA_ROSTER_NEG_TTL and a separate _letta_roster_fetched_at float precisely to work around not being able to cache it. Check whether that whole negative-TTL apparatus is still needed once the entry is the sentinel — if it is not, say so, because it is a genuine simplification rather than a relocation.
  3. Add StaleWhileRevalidateCache as a subclass for build_agent_list, with the in-flight flag as a field on a typed entry rather than a key in the value dict, and the flag cleared in a finally:. That is the Template Method: the base decides hit/miss/expiry, the subclass decides what a stale hit does.
  4. Then re-grep for a seventh. The census counts six, but it counts by shape. A read-through cache written with a different variable name is invisible to that grep and visible to grep -n "'ts': 0.0\|time.time() -" server.py.

Tests round 14 must ship

The trap in this round. Round 13 hit rule 3's second-binding failure twice, and round 14 is far more exposed to it because it is moving state, not data. A cache object constructed in server.py closes over whatever it was handed; a test that monkeypatches server.get_letta_id afterwards will not be seen unless the cache calls through the module at fetch time. The symptom is not a red test — it is a green one that took 30 seconds and talked to the live Letta API. Round 13's own notes on where this bit are in the section below.

What round 13 actually did — config out, typed landed d04b519d

−579/+192 in server.py, which fell 7,533 → 7,146. Module-level assignment went 704 → 262 lines and the ~430 lines of config literal are gone. 3,162 tests pass, up from 3,047. It is the first round to beat its own estimate (~530 predicted, 579 landed), and the reason is round 12: the wiring left behind went into a port method instead of back into the module namespace.

Where everything went

ModuleModelsReplaces
finance/reporting_categories.py ReportingCategory the four parallel REPORTING_CATEGORY_* dicts
hardware/scanners.py ScannerSpecSCANNERS
servers/registry.py ServerSpec, NamedCheckProbe, HttpProbe, TcpProbe, LogOnlyProbe SERVERS
servers/restart.py RestartCommand, RestartRegistry RESTART_HANDLERS + RESTARTABLE_KEYS
agents/registry.py LettaAgentSpec, AgentCard, VoiceOption LETTA_AGENTS, AGENT_CARDS, AGENT_VOICE_OPTIONS
finance/report_registry.py ReportMonth, FinanceReportSpec the two parallel month dicts, ROL_FINANCE_REPORTS
intake/statuses.py TerminalIntakeStatus (Literal) _TERMINAL_INTAKE_STATUSES
intake/progress.py MazdaProgressStep _MAZDA_PROGRESS_LABELS

Every legacy name stayed as a derived view — the same dict, the same list, the same per-entry key order — so not one of the ~60 consumers and not one of the ~40 tests that monkeypatch them had to move. That is what made a 579-line removal near-zero risk.

The two findings

A fix, and it was live. LETTA_AGENTS listed Shelia twice, identically. The roster is a list and build_agent_list() iterates it, so /api/agents was serving 21 tiles for 20 agents and Agent Management rendered two identical Shelia cards. Confirmed against the live dashboard before the fix and again after. AGENT_CARDS carried the same duplicate as a repeated dict key, where Python silently keeps the last — which is why the card text looked right and hid the roster bug entirely. This is number seven of the round-6 postscript's family, and the mildest: unlike the other six it was visible, and nobody had looked.
Handed on, not fixed. voice/config.py's KNOWN_AGENT_NAMES says it is “kept in sync with LETTA_AGENTS” and is not. Toyota — the receptionist the voice path routes through — has no mishear correction at all, and one minion is Suzuki Patch on the roster but Suzuki Patcher in the voice list, so a spoken “Suzuki Patcher” is corrected to a name no agent answers to. Editing it changes what whisper hears, which is a behaviour change and does not belong in a config-typing commit (rule 15). The exact drift is recorded in agents/registry.py and asserted in tests/test_agents_registry.py, so it cannot widen while it waits. Whoever takes the voice pipeline next should take this.

Everything else in the round was a defence, and rule 11 says to say so. The REPORTING_CATEGORY_* family guards a fallback path superseded at runtime by ICategoryTaxonomy — exactly as the diagnosis predicted. SERVERS, which the diagnosis suspected twice, turned out to be clean: no mistyped unit, no dangling depends_on, no unknown check name.

The shapes that earned their keep

A model that only mirrors its dict is still worth writing — that is the Pydantic section's argument and it holds. But five of round 13's models check something the dict could never express, and those are the ones to copy the technique from:

The golden test, done properly

Rule 9 asks for the old literal and the new view side by side on real live values. Round 13 did that and made it permanent: the tests read the pre-refactor literals out of git at the baseline commit, ast-evaluate them against this box's live values, and diff. A git object is immutable, so the comparison cannot decay into a restatement of the new code the way a pasted copy does — and it stays honest about the values that differ per machine (PORT, the Letta URL, two log paths). See tests/test_servers_registry.py.

Every view came back byte-identical including per-entry key order. The one intended difference is the duplicate roster entry, and the live /api/agents diff before and after the deploy shows exactly that and nothing else.

Rule 3 bit twice. Both times it was loud.

Worth reading before round 14, which is more exposed to this than round 13 was:

DIRECT_SERVICES needed no entry this round, and that is worth understanding rather than copying. Round 13 pointed routes at ports, and every port adapter resolves through server at call time — so ServiceRecorder's stubs still land. Point a route at an owning module instead and they do not. Rule 18 stands.

The ceiling, and the rise it absorbed

147/116 → 134/109. Eight names left the ladder: the seven config names the round moved (SERVERS, RESTARTABLE_KEYS, LETTA_AGENTS, ROL_FINANCES_REPORTS_MONTHS, ROL_FINANCES_REPORTS_DEFAULT_MONTH, ROL_FINANCES_REPORTS_URL_PREFIX, _rol_finance_reports_for_month) plus get_letta_id, which went with the receptionist lookup that was its only caller there.

The ceiling test was red on arrival, and that is the mechanism working. A parallel feature (e96ace55) added srv._resolve_expense_receipt_path without lowering anything, so tests/test_http_app_ports.py was failing before round 13 wrote a line. That name belongs to DocumentPort and leaves in round 20; round 13 absorbed the rise into its own reduction. If you arrive to a red ceiling, that is what it is for — find who added the name, note which port owns it, and absorb it. Do not raise the constants.

Three ports got their first methods

Each replaces a question the ladder was answering by joining globals, which is the same lesson ScannerPort.image_path() taught in round 12. Ask what the caller is trying to do:

What round 13 deliberately did not do

What round 12 actually did — ports, and the death of the free srv name landed 6c69a3f4

Kept because you will be extending this mechanism, and because one part of it is a trap that stays invisible until it bites. Removed 41 lines from server.py; the point was never the number.

What exists now

FileWhat it holds
http_app/ports.py Fourteen Protocol classes, one per collaborator group. Protocol, not ABC: the implementations are plain modules and objects, checked structurally, so a later round can swap a module-backed adapter for a real service class without touching a route. ScannerPort is populated; the other thirteen carry the names they will absorb in their docstrings and are filled in by the round that moves their code. Resist a fifteenth — a grab-bag MiscPort is srv renamed.
http_app/registry.py One frozen Ports dataclass and current_ports(), which resolves server out of sys.modules and builds a fresh bundle per call. Ports carries a field per populated port only; thirteen empty adapters built per request would answer no question. Add your field when you populate your port.
tests/test_http_app_ports.py 55 tests. Production wires the real object; the bundle is rebuilt per call and a port constructed before a rebind still sees it; all 38 dropped names are asserted absent; and the srv. count is pinned as a ceiling that only ever falls.

The three conversions

The trap round 12 leaves you. Read this one twice. tests/http_app_harness.py makes route tests inert by replacing every function on server. A route pointed at its owning module is no longer covered by that. The symptom is not a red test — it is a green one that took 30 seconds and SSHed to another box, drove a scanner, or ran whisper.

DIRECT_SERVICES in that file maps each converted module to the names it owns, and the stub is recorded under the bare name so stubbed.called('pc_metrics') reads the same on either side of the move. Every round that converts a port must add to it.

The fourteen ports, and what each still owns

Sizes are how many of the original 167 srv names each port absorbs. Eleven names get no port — they are plumbing, and they went the same way as the 47.

PortNamesOwns · populated by
ReportsPort24 Report discovery, path aliasing, status classification, the month and receipt-only queries, and the three HTML builders. The biggest port, and the one most likely to want splitting once round 16 has moved the HTML out. Rounds 16, 20, 25.
AgentsPort23 Roster, cards, voices, models, OAuth accounts, activity, health, headless runs. Round 24 — only tractable after 14.
ModelStatsPort16 Stats sources, mute overlay, Codex sync, ChatGPT provider accounts. Round 24.
ServersPort14 The Server Management tab: status, logs, restart, deploy. Round 23, after 13 moves SERVERS out.
MonitoringPort12 PC metrics, SSH roster, Win10 containers, failure classification. Every name behind it already lives in monitoring/, so round 12 pointed the routes straight at those modules — this port may never need populating at all.
DocumentPort11 Receipt lookup, supporting documents, presence checks, Excel render. Rounds 19, 20.
ExpensePort10 Manual entry, stored-expense edit and search, notes. Round 21.
IntakePort9 The recent-intake record, the halt file, scanner intake lookups. Round 15.
ScannerPort8 Populated, round 12. Scanner hardware and its diagnostics tab — small, self-contained, and a tab you can watch work.
PipelinePort8 Document processing, statement break-up, Mazda fill, reprocess. Rounds 17, 18.
VoiceNotesPort8 Voice upload, speech synthesis, the receptionist strategy, note commands. Round 12 pointed the routes at voice/ directly; the port populates when the voice pipeline gets a composition root.
TerminalPort6 PTY spawn and reap, WebSocket framing — already terminal/pty_session.py and http_app/websocket.py. Round 12 pointed terminal_ws.py at both directly.
CategoryPort5 Recategorize, undo, vendor review. Already behind finance/recategorize.py.
MazdaPort2 The mode switch. Round 11 reduced it to two verbs — it is the shape the other thirteen are aiming at.
no port11 HERE, REPO_ROOT, LETTA_BASE_URL, ValidationError, _load_json/_append_json/_clear_json, and the four Claude log files and locks. None of these is a service. Import them from the module that owns them — and note that a route holding a threading.Lock reached across a module boundary is a bug waiting to be written, not an interface. Three of the four Claude log locks are still reached through srv; they are the ugliest thing left on this list.

Pydantic: many small models, deliberately

Rule 10 is amended. It used to read “reach for Pydantic where a wrong value is silent, not everywhere.” That was right for eleven rounds of careful, narrow extraction and it is wrong for this phase. The new reading: every payload that crosses into the browser or onto disk gets a model, and every config literal becomes a typed collection. The judgement that rule 10 protected does not disappear — it moves to which field is required, which is where the thinking always was.

The justification is mechanical, not stylistic. contracts.StrictModel is strict=True, extra='forbid', frozen=True. A model that merely mirrors the dict it replaces is therefore not decoration: it turns all three of the silent failures this plan keeps finding — a coerced value, an unexpected key, a mutation in flight — into exceptions at the boundary. Mirroring is the safety. That is what makes writing thirty of them defensible.

Against 152 return { boundaries and about 430 lines of config literal, the dashboard currently has 35 StrictModel subclasses. It also has 33 plain BaseModel subclasses — a second, non-strict, non-frozen, extra-permitting base sitting alongside the strict one. Audit those as you pass them; most should derive from StrictModel, and the ones that genuinely cannot should say why in a docstring.

Config models — round 13's worklist done, d04b519d

All thirteen shipped, in seven modules. What follows is the worklist as it was written, with what each one actually turned out to be — the predictions were mostly right, and the two places they were wrong are the instructive ones.

ModelReplacesPredicted failure What it turned out to be
ReportingCategory 4 parallel REPORTING_CATEGORY_* dicts (58 lines) A category in one dict and missing from another — a report row with no colour, served as if styled. Defence, as predicted: the maps guard a fallback path superseded by ICategoryTaxonomy. Also caught what the plan missed — the ancestor map is not derivable from db_id alone, so the model carries ancestor_ids and validates that a bucket's own id rolls up to itself.
ServerSpec + four probes SERVERS (159) A mistyped systemd unit or host: a Restart button that reports success and does nothing. Clean. No mistyped unit, no dangling depends_on, no unknown check name. The union is over the active probe, not over log_file — see why.
AgentCardAGENT_CARDS (136) A card missing a key renders as a blank tile rather than an error. Clean as copy, but it carried a repeated dict key (Shelia) that Python resolved silently — which is what hid the live roster bug next door.
ScannerSpecSCANNERS (23) A wrong device or namelike scans the wrong hardware, or nothing, quietly. Defence, and the model went further than typing: it cross-checks that the matchers appear in the device string.
RestartCommand + RestartRegistry RESTART_HANDLERS + RESTARTABLE_KEYS (24) A key in one and not the other. Those two could never disagree. The real gap was that neither was tied to SERVERS — so the registry checks coverage of the tiles instead.
LettaAgentSpecLETTA_AGENTS (22) A malformed roster entry becomes unknown-<name> downstream instead of raising. A live fix, and the round's headline. Not the predicted failure at all — a duplicated entry serving a duplicate agent tile. The finding.
VoiceOptionAGENT_VOICE_OPTIONS (18) An unknown voice id reaches edge-tts and fails at speech time, on a background thread. Defence. All sixteen well-formed. Led to the voice-config drift finding next door.
FinanceReportSpecROL_FINANCE_REPORTS (16) A dir that does not exist becomes an empty tab. Defence. The model checks the dir is a bare folder name, since it is joined to the month's base dir.
ReportMonth ROL_FINANCES_REPORTS_MONTHS + ROL_FINANCES_MONTH_RANGES (12) Two more parallel dicts on the same four keys. A month in one and not the other is a tab whose status query silently returns nothing. Defence, plus the stronger invariant the plan did not ask for: the range must be the month the key names.
MazdaProgressStep_MAZDA_PROGRESS_LABELS (11) An unlabelled stage renders as a blank progress line. Defence, and it found the sharper risk: the consumer indexes a parallel statuses list by position.
TerminalIntakeStatus (Literal) _TERMINAL_INTAKE_STATUSES (10) Round 11 flagged this as the natural first Literal of the intake block. Done. It is a vocabulary now, with the reason it matters written down: an unknown status is dropped, leaving the document on processing forever.

Three of the worklist did not move, and the reasons are worth keeping. HealthCheckSpec became CheckName, a Literal in servers/registry.py, because HEALTH_CHECKS maps names to functions defined in server.py — the vocabulary can leave, the bindings cannot until round 23. A test asserts the two agree in both directions. ModelOption and OAuthProviderAccount were left for round 24, which owns Model Stats: AGENT_MODEL_OPTIONS is already consumed through AgentModelOptionsService, so typing it away from that service would model the list rather than the boundary.

The 33 plain BaseModel subclasses are still unaudited. Round 13 added none — all thirteen of its models derive from StrictModel, with exactly one documented relaxation (RestartCommand sets arbitrary_types_allowed, because a handler is a callable; it stays frozen and extra-forbidding). The dashboard now has 48 StrictModel subclasses, up from 35.

Payload models — a first tranche from the 152

Take these as each round reaches them; do not do a payload-typing round of its own, because a model written away from the code that builds it just mirrors today's bug.

RoundModels
15–16 RecentIntakeRow, IntakeHaltRecord, ScannerIntakeRef, MazdaProgress
17–18 StatementPreflightPayload, StatementRecord, PipelineResult, ProcessedDocument
19–20 DocumentReference, SupportingDocumentView, ReceiptLookupResult, ReceiptOnlyRow, MonthStatus, RecentScanRow
21 ManualEntryResult, ExpenseEditResult, StoredExpenseEvent
22–23 ScannerDiagnostic, ScannerRuntimeStatus, ServerHealthPayload, DeployResult
24–25 AgentActivity, AgentHealth, WeeklyRemaining, ModelStatsAgentRow, CodeStatus, ReportAttention

The discipline that does not relax: keep the golden-payload test against the prior literals, error strings and key order included, and where you can, run old and new side by side against real live values. Rounds 9, 10 and 11 all did it and all three caught something.

JavaScript: many small interfaces

js/ is 35,400 lines and it is not in bad shape structurally — there are already 39 files in js/abstract/, 43 in js/implementation/ and 76 test files, and the live entry point js/dashboard-boot.js is 157 lines. The convention is established and it works. The problem is granularity: the interfaces are too few and too big, and three implementation files have grown into the same shape as server.py.

The rule

One interface per collaborator, not one per file. An interface exists so a caller can depend on the part it uses. A 470-line interface serving one 1,365-line implementation is not an abstraction, it is a table of contents — and it is the Interface Segregation Principle violated in exactly the way server.py violates it in Python. If an interface's consumers each use a different third of it, it is three interfaces.

Where it stands

FileLinesInterfaceVerdict
implementation/rol-finance-reports-controller.js 1,300none The worst file in js/. One class, 20 constructor parameters, 9 endpoints, and at least eight responsibilities: month tabs, month detail, recent report, scanner tabs, iframe loading, a picker dialog, a headless Mazda UI, and polling.
implementation/manual-entry-form.js 1,365 manual-entry.interface.js (470) The interface is genuinely good — real typedefs, pure decision logic, no DOM, mirrors finance/manual_entry.py's rules. It is simply doing five jobs. Split the interface, and the form follows it.
implementation/detail-renderers.js 1,464 detail-renderer.interface.js (44) Best-shaped of the three: 15 exports behind a small, correct interface. It is fifteen renderers in one file. One file each, plus a DetailRendererRegistry, and the interface does not change at all.
dashboard-boot-baseline.js 3,101— Referenced only by dashboard-baseline.html. Dead weight in every grep and every search result. Confirm nothing serves it, then delete both — or, if it is deliberately frozen, say so in a header comment so the next agent stops re-reading it.

The named splits

rol-finance-reports-controller → 5 interfaces

ReportFetcher · MonthNavigator · ReportTableView · CategoryPickerDialog · ReprocessCommand

Take CategoryPickerDialog first — _ensurePickerDialog through _closePicker is ~95 self-contained lines with its own DOM subtree, and it is the one piece with an obvious test.

manual-entry.interface → 5 interfaces

FieldValidation (already exists as its own file) · VendorResolution · AmountEntry · SaveRequestBuilder · StoredFindingReader

The typedefs already cluster this way — ManualEntryValidation, ManualEntryPrefill, StoredFindingRow are three separate concerns sharing one file.

detail-renderers → 15 files + a registry

Mechanical, low-risk, and it is the one that makes the js/tests/detail-renderers.test.js (615 lines) legible again. Registry replaces whatever switch currently picks a renderer.

Every new interface gets an interface test

The convention already exists — js/tests/manual-entry.interface.test.js (616 lines) and js/tests/interface-workspace.test.js. A new *.interface.js without a matching *.interface.test.js is unfinished.

A cross-language finding worth acting on early. RolFinanceReportsController's constructor hardcodes, as default arguments, the four month keys (jan-2025…apr-2025), the two scanner keys (window, freezer) and the Mazda agent id. Python holds the same four months in two parallel dicts (ROL_FINANCES_REPORTS_MONTHS and ROL_FINANCES_MONTH_RANGES), the same two scanners in SCANNERS, and the same agent id in MAZDA_AGENT_ID. That is three copies of the month list and two of the scanner list, across two languages, none of them checked. Round 13 makes the Python side one typed collection; the JS side should then take its list from an endpoint rather than a default argument. Until it does, one plan rule applies unchanged: one destination, one definition.

Rule 15 supersedes the old “stay out of js/”

Rule 15 used to say: stay out of js/ unless the bug is there, because a concurrent dashboard-boot.js refactor runs in parallel. That refactor has landed — the live boot file is 157 lines and the module tree under js/boot/ is real. The rule becomes: touch js/ deliberately, one interface at a time, and still run git status js/ first, because another agent may be mid-edit. Do not mix a JS split and a Python extraction in the same commit; they fail differently and you want to be able to revert one.

How to rank a cluster — still useful, now secondary

The residual-coupling script below has picked every round since round 8 and it is still the right tool for ordering work inside a round. It is no longer how rounds are chosen: the burn-down is fixed and ordered by dependency, not by weekly measurement. Use the script to decide what travels with a block once you have committed to that block, and to check that a block really is as self-contained as this plan claims.

The right question it answers: which names stay behind in server.py after the move? A name referenced only by the cluster travels with it and costs nothing.

import ast
from collections import defaultdict
src = open('server.py').read(); tree = ast.parse(src)
funcs = {n.name: n for n in tree.body
         if isinstance(n, (ast.FunctionDef, ast.AsyncFunctionDef))}
top = set(funcs)
for n in tree.body:
    if isinstance(n, ast.Assign):
        for t in n.targets:
            if isinstance(t, ast.Name): top.add(t.id)
refs = defaultdict(set)          # name -> top-level funcs that reference it
for fname, node in funcs.items():
    for sub in ast.walk(node):
        if isinstance(sub, ast.Name) and isinstance(sub.ctx, ast.Load) \
                and sub.id in top:
            refs[sub.id].add(fname)
def deps(name):
    n = funcs.get(name)
    return set() if n is None else {
        s.id for s in ast.walk(n)
        if isinstance(s, ast.Name) and isinstance(s.ctx, ast.Load) and s.id in top}
names = {...}                    # the cluster you are considering
while True:                      # pull in what only this cluster uses
    ext = set().union(*(deps(n) for n in names)) - names
    travels = {x for x in ext if not (refs[x] - names)}
    if not travels: break
    names |= travels
stays = sorted(set().union(*(deps(n) for n in names)) - names)
lines = sum(funcs[n].end_lineno - funcs[n].lineno + 1 for n in names if n in funcs)
print(lines, 'lines |', len(stays), 'stay behind:', stays)
print('cluster members:', sorted(names))

Caveats, all of them earned

  1. The script counts function bodies, so module-level constants that travel are free lines it does not credit — and the wrappers, composition root and explanatory comments left behind cost lines back. Round 11 measured 153 and removed 107; round 10 measured 210 and removed 214; round 9 measured 285 and removed 374 (a 70-line config literal the script cannot see); round 8 measured 183 and removed 152. Assume ~70% of the measured figure for a cluster with no constants — and expect that ratio to improve after round 12, because the wiring left behind moves into the container instead of into the namespace.
  2. It cannot see http_app/. It parses server.py and nothing else, so every srv.<name> call in the route ladders is invisible to it. Seven such names in round 8, nine in round 9, zero in round 10, two in round 11 — and round 11's own table said zero, because that column is a grep from the round before. Always re-grep. Since round 12 the suite catches half of this: tests/test_http_app_ports.py fails if the srv. count moves in either direction. It still cannot catch a name you moved out from under a live srv. caller, because tests/http_app_harness.py stubs the service layer and those tests pass against a name that no longer exists.
  3. A route reaching into a private name is a design signal. Round 9's four became two verbs and one was fixing a real race; round 10's poke at _chatgpt_failover['last_note'] became last_failover_note(); round 11's two pokes at _MAZDA_MODE_SERVICE became mazda_mode_status() and set_mazda_mode(). There are 26 private names left in the 169. Ask what the caller is trying to do before you give it a port method that just re-exposes the private thing.
  4. It cannot tell you a name is mis-clustered. Round 10's record_agent_send_error would have dragged three cache globals along and broken a route. Round 11's _watch_intake_for_problems would have dragged the whole Trainer escalation service into a dispatch module — it became an injected port instead. Round 8's “Document Vision health” 83 lines with zero stays was a mirage over an already-extracted registry. Grep every name you are about to move, and ask which cluster it would belong to if both were already extracted.
  5. Cluster membership is a judgement. Widen or narrow the set and the numbers move.
  6. The burn-down is a plan, not a measurement. Its per-round figures come from the AST census, so they inherit caveat 1. If a round lands 30% under, that is normal; if two consecutive rounds do, stop and re-read the diagnosis before taking a third.

Standing rules for every round

Rules 1–11 were paid for in production incidents and are unchanged except where noted. Rules 12–17 are new for the sub-2,000 phase.

  1. Extract lowest-residual-coupling first, measured with the script above, not by eye — within the round the burn-down assigns you. Then grep every name in your set: the script cannot tell you a name is mis-clustered, and it will hand you names that belong to somebody else's cache or somebody else's service.
  2. Grep http_app/ before dropping any name, even when a table says zero — that column is always a grep from the previous round. The HTTP tests stub the service layer, so they will not catch a missing name either.
  3. Point new tests at the owning module, never at server. A re-export is a second binding; the moved code closes over its own module global, so monkeypatch.setattr(server, 'X', …) isolates nothing while looking exactly like it does. This has now bitten eight rounds running. The corollary: when you move code, the tests that break are the honest ones — the dangerous ones are those that keep passing.
  4. Inject what stays behind rather than importing it, and resolve the bundle per call. Add a test that production wires the real collaborators, and one that proves the late binding. A background thread needs extra care: take the sweep as a callable so the bundle is rebuilt per iteration instead of frozen at boot (round 10's poll_loop). A wrapper whose body is one collaborator call is not a smell — it is how a neighbouring cluster's boundary gets drawn without moving it (round 11's watch_intake).
  5. One destination, one definition. A dependency that already has its own module (LETTA_BASE_URL, LETTA_DOCKER_HOST, classify_failure) is imported, not injected. So is a vocabulary: if the destination module already spells out the values you are about to re-declare, annotate against its alias. As of this round the rule crosses the wire too — see the JS finding about the month and scanner lists.
  6. Import back only what is still called, and re-check the previous round's re-exports, which expire when their last caller leaves. Round 10 dropped round 9's _probe_codex_usage; round 11 dropped DispatchEvidence and InvalidExecutionMode. Assert the dropped names are absent so dead re-exports cannot creep back. Round 12 retired the bulk of this rule, which was its point: 38 names left and tests/test_http_app_ports.py asserts each one absent by name. What survives is the judgement call — when your round's last caller leaves, check whether server.py still calls it before you keep the import.
  7. Deploy and verify against the real thing. Committing is not deploying, and a 200 is not a verification. For a backend-only change, call the moved code on the live box and read the output. If a path cannot be fired safely, say which one and why rather than firing it or pretending it was checked — and find the reversible neighbour that exercises the same machinery (round 10 re-parked an identical standby bundle; round 11 ran the real merge against a copy of the live pointer file and drove the write route with the value it already held).
  8. A fake must be pinned to the real thing it fakes. Where a dashboard module fakes a rol_finances object, add one test that runs the path unpatched against the real one. (See the postscript.)
  9. Every model dumps to a plain dict so browser-facing JSON stays byte-identical, and keep a golden-payload test against the prior literals — error strings and key order included. Better still, run the old code and the new one side by side against the real live values; it is a five-line script and it converts “should be identical” into “is identical”.
  10. Model the boundary; think hard about which field is required. Amended this round — the old wording ("not everywhere") is superseded by the Pydantic section. What survives intact: require the field whose absence disables a guard; leave optional the field whose absence a recovery path is built to handle; check the predicted defect is actually reachable; write the test that proves it as the old code, inline, so it keeps being true. And remember min_length=1 is not “not blank” when the consumer .strip()s.
  11. When a round's model defends rather than fixes, say so. Round 10 found a live defect; round 11 found five silent-but-currently- unreachable ones. Both are worth shipping, but a plan that reads them the same way teaches the next shift to overclaim. Write down which it was.
  12. An error message is an input to something. Before shipping a new failure string, run it through classify_failure() (or whatever downstream reads it) and assert the label, not the text. Also assert the strings that should keep their label still do — scrubbing that goes too far turns a real rate limit into a generic error.
  13. New — name the pattern in the commit message. “Extract X out of server.py” described eleven rounds honestly. From here the commit says which of the seven you applied, because the next shift needs to know whether you moved code or changed shape. A round that changes shape and removes no lines is a good round; say so.
  14. Never add a srv. reference. Round 12's value is destroyed by round 15 adding two. If a route needs something new, it needs a port method — or, if the thing already has an owning module, an import of that module. Enforced since round 12: SRV_SITE_CEILING / SRV_NAME_CEILING in tests/test_http_app_ports.py fail the build if the count rises, and fail it again if the count falls without someone lowering the constants. Lower them; a ceiling nobody moves becomes a lie the next round inherits.
  15. New — touch js/ deliberately, one interface at a time. Supersedes the old "stay out of js/". The parallel dashboard-boot.js refactor has landed. Still run git status js/ first — another agent may be mid-edit — and never mix a JS split and a Python extraction in one commit.
  16. New — no new module-level mutable state in server.py. Not a cache, not a lock, not a registry dict. There are 20 locks and 6 hand-rolled caches to remove; adding a twenty-first is how the file grew in the first place. State belongs to an object in the module that owns it.
  17. New — if a function is over 60 lines, it is a strategy, not a function. The 25 functions over 50 lines hold 28% of the file, and every one of them is a switch over a document kind, a scanner, a report type or a status. Do not move a 171-line function into a new module and call it extracted — that just relocates the problem and makes the next round's diff unreadable.
  18. New from round 12 — when you change how a route resolves a name, change how the harness stubs it. tests/http_app_harness.py makes route tests inert by replacing every function on server. Point a route at an owning module and that stub no longer lands, and the symptom is not a red test — it is a green one that took 30 seconds and talked to another machine. Add the module to DIRECT_SERVICES in the same commit. The wider rule: a safety net is aimed at a specific patch target, so moving the target is a change to the net, whether or not anything turns red.
  19. New from round 12 — git status again before you commit. Two agents write to this tree. Round 12 had two of its unfinished files swept into an unrelated commit by a concurrent git add -A, and wc -l server.py moved by ~170 lines mid-round for reasons that had nothing to do with the refactor. Stage deliberately, and report your round by its diff, not by the file size.

Postscript from round 6 — still the reason for rules 7, 8 and 10

finance/vendor_lookup.py called VendorCategoryLookup.guess_vendor_key, a method rol_finances has never defined — and the test that shipped in the same commit monkeypatched a fake lookup object defining that same imaginary method. The suite stayed green while every production caller raised AttributeError, and the route swallowed it into a 200 page reading “Scanner Report build error”, so monitoring saw success too. Fixed in 603d2d78.

Round 8's docker ps bug was the same animal in a third costume: a failure whose observable signature is identical to success. Round 9's classifiers were the fourth — a payload nobody could read, reported as an account with headroom. Round 10's is the fifth and the worst: a parked credential nobody could read, treated as a wildcard that matched every account, and then written back. Round 11's is the sixth, the mildest and the most instructive: a status record that failed to land, whose only signature was a document sitting on processing forever — a scanned receipt that looks, on the page, exactly like one still being worked on.

All six were invisible to the suite. All six needed someone to run the real thing and look at what came back.

Number seven arrived in round 13, and it was not either of the two this postscript predicted. The mistyped-SERVERS-entry and the missing-REPORTING_CATEGORY_CLASS-key were both hypothetical; typing them found nothing wrong with either. What was actually in the tree was a duplicated roster entry serving a duplicate agent tile — and it breaks this family's pattern in the one way worth noting: it was not invisible. It was on the screen, on the live dashboard, for anyone who counted the cards. Six failures whose signature was identical to success, and then one whose signature was plainly wrong that nobody had looked at. Both are the same lesson from opposite ends: run the real thing and read what came back.

The two predicted failures are unfound rather than disproven. The validators for both now exist, so if either is ever written it fails at import instead of in a browser.

Session log — rounds so far

916dd739 Extract the HTTP layer, hardware I/O and stats extractors out of server.py
6e3ddc96 Type the Model Stats, PC Monitor and terminal seams with Pydantic
9b82f01b Name the health-check contract, and move the two biggest checks out of server.py
6d7606d9 Extract the health-poller cache and vendor-review dialog backend out of server.py
1351c3cc Extract the Server Management clocks and log-file reading out of server.py
5f7d2059 Extract the category write and its undo out of server.py
9552ce0b Extract the Win10 box and the three agent tabs out of server.py
5ae369aa Make the Win10 container probe survive the remote shell
47778b3f Extract the SSH connection checks and the provider quota probes out of server.py
5b1fc498 Give a refused usage payload a noun in its error text
039ad71b Extract the ChatGPT provider auto-failover state machine out of server.py
78c8b9fa Extract the Mazda dispatch fork and the execution mode out of server.py
6c69a3f4 Facade: put ports between http_app and server.py (round 12)
d04b519d Registry: config out of server.py, typed (round 13)

Follow-ups that removed no lines from server.py but which the next shift inherits: dbecf4fc (plan update, round 7), e28bd5fd (re-fetch the plan when its tab opens), b25955d5 (correct the plan's own stale numbers), b3eba0e1 (stop the plan recording a hash that is stale on arrival).

Every module created so far

ModuleWhat lives there
model_stats/sources, windows, usage_history (leak detection), last_good, reader, assignments, extractors
monitoring/pc_metrics.pyPcMonitor, PcMetric (derives alert from level)
monitoring/server_lifecycle.pystarting window + down/stale clock, ServerStatus
monitoring/log_files.pytail/age/mtime-probe, server_log_rows
monitoring/win10_node.pythe Win10 box's reachability, containers, dockerd and the two recovery buttons; Win10CacheEntry, _TtlCache — round 14 promotes _TtlCache to caching.py
monitoring/ssh_checks.pythe tailnet roster, the ssh and tailscale probes, the debounced health cache, the per-connection log tail, and the Test button
monitoring/provider_usage.pythe zero-token quota pipeline; CodexUsage, ClaudeUsage, UsagePayloadError, shape_detail, PROVIDER_USAGE_PROBES
monitoring/chatgpt_failover.pythe account-swap state machine: cooldown gate, standby verdict, token heal, swap, 90s sweep; StandbyCredentials, last_failover_note
intake/mazda_dispatch.pythe fork between Mazda's LLM turn and a human's inbox: the mode branch, the scan dispatch, the late-delivery probe, and the record written when it declines; IntakeOutcome, Collaborators, HUMAN_ONLY_MODE_STAGE_MESSAGE
intake/mazda_mode.pythe Automatic/Semi-Automatic vocabulary, the operator's switch and its store, and the MAZDA_DECISION_MODE parser: MazdaMode, ExecutionModeConfig, resolve_execution_mode, MazdaModeService, MazdaModeState
intake/ (rest)dispatch_evidence.py, scan_message.py, trainer_contracts.py, trainer_escalation.py, trainer_notifier.py, trainer_recovery.py
http_app/the whole HTTP layer: transport, the two route ladders, terminal WS, runtime, static files, and srv, now being retired. ports.py holds the fourteen Protocols the ladders may depend on; registry.py holds current_ports(), which resolves through server per call and is the only module besides services.py allowed to name it.
terminal/pty_session.pypty spawn, process-group reap
letta_code/runner.pyheadless letta-code invocation + prompt validation
health/probe.py (ProbeResult), failures.py, frita.py, document_vision.py, poller.py
agents/letta_gateway.py, urllib_letta_gateway.py, model_options.py, message_views.py
finance/vendor_review.pylist_vendor_keys, list_pending_vendor_review, set_receipt_vendor, PendingVendorReviewRow
finance/recategorize.pythe category write, its undo, the report-row repaint, the undo journal's composition root, ReportRowClass, RecategorizeRequest
hosts.py, paths.py, letta_ids.pyshared constants that previously had 2–4 independent env-read/fallback copies. LETTA_DOCKER_HOST reached four in round 10.
finance/reporting_categories.py round 13. ReportingCategory and the thirteen reporting buckets; the four legacy REPORTING_CATEGORY_* dicts are derived views of it. NOT the runtime source — ICategoryTaxonomy is; these guard the fallback. category_taxonomy_seed.py is a fifth copy and is deliberately not derived from this.
hardware/scanners.py round 13. ScannerSpec; cross-checks that namelike/driver_match describe the device the script drives. SCANNERS is a derived dict view — the sixteen call sites still read dicts, and round 22 converts them with the hardware.
servers/registry.py round 13. ServerSpec with a discriminated ServerProbe union (NamedCheckProbe, HttpProbe, TcpProbe, LogOnlyProbe), CheckName as the check vocabulary, and build_server_specs() — a factory, because four entries interpolate values the composition root owns.
servers/restart.py round 13. RestartCommand + RestartRegistry (Command). Owns dispatch and the check that every Server Management tile has a Restart button. The handlers themselves stay in server.py until round 23.
agents/registry.py round 13. LettaAgentSpec, AgentCard, VoiceOption, and the roster/cards/voice catalogue. Refuses a roster with a repeated name or Letta id — which is how the live duplicate-Shelia tile was found. Also records the voice/config.py drift.
finance/report_registry.py round 13. ReportMonth (the two parallel month dicts, with the range validated against the key) and FinanceReportSpec. MONTH_KEYS is what the JS controller should read instead of its hardcoded default arguments.
intake/statuses.py, intake/progress.py round 13. TerminalIntakeStatus (the intake block's first Literal, flagged by round 11) and MazdaProgressStep, which pins position == step number because the consumer indexes by position.
contracts.pyStrictModel — the one strict, frozen, extra-forbidding base every boundary model derives from. 48 subclasses after round 13, up from 35. 33 plain BaseModel subclasses still bypass it; audit them as you pass.
ssh_gateway.pydead — a second, subtly wrong SSH probe nothing imports. ConfiguredIdentityStrategy returns the first identity file that exists, which is the bug ssh_test's docstring exists to describe: a key can exist and no longer be authorized. Left in place (another agent's parallel work), but do not build on it, and port the fallback chain in first if a future round wants the adapter shape. tests/test_ssh_checks.py names the trap.

Where Pydantic mattered, cumulatively

Round 13's thirteen, in one line each. ReportingCategory — a bucket's own id must roll up to itself, so the picker cannot write a category the report then attributes elsewhere. ScannerSpec — the diagnostics matchers must appear in the device string the scan script drives. ServerSpec + its four probes — exactly one active probe, because server_health() resolves check before health_url and an entry with both advertised a URL it never pinged. RestartCommand — a key and its handler stop being two facts, and the registry must cover every tile. LettaAgentSpec — no repeated name or Letta id, which is what caught the live duplicate tile. AgentCard — a card cannot be half-filled. VoiceOption — a voice id is checked where it is chosen, not on a background thread at speech time. ReportMonth — the calendar range must be the month the key names. FinanceReportSpec — a report dir is a bare folder name, not a path. MazdaProgressStep — position equals step number, because the consumer indexes by position. TerminalIntakeStatus — one Literal for the vocabulary whose misses are silently dropped.

Round 12 added no model, and that is the honest report (rule 11). It changed shape: fourteen Protocols and a frozen Ports bundle. The one thing it typed by accident is worth noting anyway — ScannerPort.image_path() replaced a route that reached for srv.SCANNERS[key]['output'], so the browser side of the scanner config now has exactly one reader. That is what makes ScannerSpec cheap in round 13.

Next, per the worklist: round 14 adds no config model at all — it replaces six hand-rolled caches with one class, and its typed entry (CacheEntry, promoted from Win10CacheEntry) is the model that matters, for exactly the reason Win10CacheEntry mattered: the entry, not the value, is the presence sentinel, so a cache whose legal answers include None can actually cache one. Five of the six get that wrong today.

How this plan reaches the dashboard

It is already wired in: dashboard.html has a #dashboard-refactor-plan-frame iframe pointing at /notes_plans_handoffs/dashboard_refactor_plan.html, behind the Dashboard Refactor tab under Project Plans. Live at https://desktop-2obsqmc.tailb8fc54.ts.net/.

A trap that looks exactly like a failed deploy: a tab switch only toggles a CSS class, so the iframe is fetched once, when the dashboard page itself loads. A dashboard left open all day keeps showing whichever revision it pulled at load time, however many times you redeploy. The frame carries data-refresh-on-show, and activateView in js/boot/view-navigator.js re-fetches marked frames every time their tab is opened. If a plan still looks stale, reload the dashboard page before suspecting the deploy — and confirm with curl -s <live>/notes_plans_handoffs/<plan>.html | grep <marker>, which asks the server rather than the browser.