Living Plan · Round 13 landed 2026-08-26 · Round 14 brief
The target has moved. dashboard/server.py is no longer being
halved — it is being reduced to a composition root of
under 2,000 lines, and designed to land near 700. The
method changes with it: stop carving off whichever cluster measures best
this week, and instead dismantle the three structures that are holding
7,500 lines up. Many small Pydantic models. Many small JavaScript
interfaces. Patterns with names.
Two of the three are down. Round 12 put ports between the
routes and server, so moving code costs a port method instead
of a permanent re-export. Round 13 spent that: every hand-maintained config
literal is gone, module-level assignment fell from 704 lines to
262, and the roster it typed turned out to be serving a duplicate
agent tile to the live dashboard. What is left of the three is the
six hand-rolled caches — round 14, and the
thing the whole agent block is pinned behind.
You are picking up round 14 — one cache class, six instances. Do these six things before writing any code.
git status again when you are
about to commit. Round 12 was mid-flight when a concurrent agent
ran git add -A and swept two of its unfinished files into an
unrelated commit. Nothing was lost, but the tree is genuinely shared and
“I checked at the start” is not the same as “it is still
true.”DIRECT_SERVICES in
tests/http_app_harness.py. Round 13 needed no entry there,
because it routed everything through ports rather than at owning
modules — but it hit the same trap twice from the other side, and
its notes say where.server.py actually is when you arrive. Every number
on this page is from d04b519d and every number on this page
will be stale for you — a second agent ships features into this file
while you work, and round 12 measured a 29-line growth across a
round that removed 41.The halfway target was reachable and the plan was on course for it. That is exactly the problem: 5,700 lines of procedural code with no classes in it is not a better building than 11,407 lines of the same. It is the same building with a third of the rooms in the garage. The rounds were getting slower and smaller — round 11 measured 153 lines and removed 107 — because the cheap, cohesive clusters are gone and what remains is held in place by structure, not by size.
Round 11's own recommendation was to stop taking small slices and spend a round designing a seam. That instinct was right and the scope was too modest. The retarget makes it explicit:
“Which cluster can I move this week with the fewest names left behind?”
Answer quality degrades every round by construction: you are always picking from a set that gets worse.
“What is server.py for, and what has to be
true before everything that isn't that can leave?”
Answer: it is a composition root. About 700 lines. Everything else is someone's module.
Under 2,000 is the ceiling, not the design. The design is itemised below and comes to roughly 700. The 1,300 lines of daylight between them is the slack this plan has learned, over eleven rounds, that it always needs.
Measured on d04b519d. Reproduce it from dashboard/
with the AST script in How to rank a cluster plus
these four one-liners:
grep -oh 'srv\.[A-Za-z_][A-Za-z0-9_]*' http_app/*.py | sort -u | wc -l # 111 (109 + 2 docstring artefacts)
grep -c 'return {' server.py # 162
grep -o '\.get(' server.py | wc -l # 473
grep -c 'with .*_lock' server.py # 36
Report your round by its diff, not by wc -l.
Round 13 landed −579/+192 in server.py and the file moved
7,533 → 7,146 — those do not reconcile, and they are not
meant to. Two agents write to this tree: round 12 removed 41 lines while a
parallel feature added ~170 on the same days, and round 13 arrived to find
the srv. ceiling test already red for the same reason.
classes defined in 7,147 lines. Not “few”. Zero. Every
piece of state in this file is a module global and every piece of
behaviour is a top-level def. Rounds 12 and 13 added
classes to the HTTP side and to the new registry modules; this
file still has none of its own.
top-level functions, 5,396 lines of bodies. 25 of them are over 50 lines and account for 2,093 — 28% of the file lives in 10% of the functions. Round 13 moved data, so this number did not move at all. Rounds 15–21 are what shift it.
distinct names http_app/ still reaches into
server for, across 134 call sites — down from
167 / 220 before round 12 and 117 / 148 before round 13.
22 of what is left are private. What round 13
did.
lines of module-level assignment, down from 704. The
~430 lines of config literal are gone —
SERVERS, AGENT_CARDS, LETTA_AGENTS,
SCANNERS, the four REPORTING_CATEGORY_* dicts
and six smaller ones are now typed collections in seven modules.
Round 13.
module-level threading.Lock() objects, 36
with …_lock: blocks, 7 hand-rolled TTL constants,
and six independent hand-written read-through caches
— while monitoring/win10_node.py already contains a
tested _TtlCache class.
return { statements — untyped payload boundaries,
most of them going straight to the browser — against 473
.get() calls, each one a place a missing key becomes a
default instead of an error. The dashboard now has 48
StrictModel subclasses — round 13 added 13.
Eleven rounds of extraction did not shift any of these, because none of them is a cluster. Each one has a name in the literature and a standard remedy. Two are now done; the third is round 14.
srv is a Service Locator anti-pattern
mechanism built, round 12
http_app/services.py exposes srv, a PEP 562 module
__getattr__ onto server. The two route ladders
reached through it for 167 distinct names across 220 call
sites, 26 of them private (srv._fetch_month_status,
srv._load_json, srv._voice_log_lock — a
lock, reached across a module boundary, by an HTTP route).
This is why every round paid a re-export tax. Move a function out and the
route still says srv.the_name, so the name has to stay behind
as an import. 47 names in server.py existed for no other
reason, and they would never have expired on their own, because their
caller was the one thing not moving.
The remedy: Facade + Registry.
Round 12 built it —
http_app/ports.py and http_app/registry.py —
converted all 47 free names, and converted ScannerPort end to
end as the worked example. Round 13 took the ceiling to 109 names
/ 134 sites, and tests/test_http_app_ports.py pins it
as a ceiling that only ever falls.
What is left is the interesting half. The remaining 109 are real services, not re-exports, and they leave one at a time as their round populates its port. Your job includes lowering that ceiling by the names your block owns — and the mechanism is working: round 13 caught a parallel feature that had added a name without lowering anything, because the ceiling test was red before a line of round-13 code was written.
About 430 lines were hand-maintained data literals: not logic, never typed,
and several of them parallel — multiple dicts keyed alike that
had to agree, with nothing checking that they did. They are now seven typed
registry modules, and round 13 is the record of
what that found. The diagnosis below is kept as written, because two of its
predictions are worth reading against what actually happened: the
REPORTING_CATEGORY_* family turned out to be a defence
exactly as predicted, and SERVERS turned out to be clean —
while LETTA_AGENTS, which this section never suspected, was
serving a duplicate agent tile to the live dashboard.
REPORTING_CATEGORY_DB_MAP, REPORTING_CATEGORY_CLASS
and REPORTING_CATEGORY_STYLE are three dicts keyed by the same
thirteen category names, and REPORTING_CATEGORY_ANCESTOR_MAP is
a fourth keyed by the ids the first one hands out. Adding a category to one
and forgetting another produces a report row with no CSS class or no colour
— rendered, served, and indistinguishable from a styling choice. One
ReportingCategory model with a
name / db_id / css_class / background / font shape, one list of
thirteen, and the four dicts become derived views that cannot disagree.
58 lines become about 20, and a whole class of silent defect stops being
expressible. Read the comment above
REPORTING_CATEGORY_ANCESTOR_MAP before you start: these
four are now the offline seed for LEGACY_TAXONOMY, superseded
at runtime by ICategoryTaxonomy. That makes them safer to
restructure, and it also means the validator you add is guarding a fallback
path — say so, per rule 11.
SERVERS (159 lines) and AGENT_CARDS (136) are the
same shape of risk at four times the size: a mistyped systemd unit name in
SERVERS is a Restart button that reports success and does
nothing — the exact failure signature the
round-6 postscript exists to warn about.
_agent_list, _agent_activity,
_agent_health, _letta_id,
_weekly_remaining and _model_stats_agents each
carry a {'value': None, 'ts': 0.0} dict, a
threading.Lock(), and a *_CACHE_TTL constant, with
the read-through logic re-typed inline at every call site.
build_agent_list adds a stale-while-revalidate background
refresh to its copy and stores the in-flight flag in the same untyped dict.
monitoring/win10_node.py already has the class this wants
— _TtlCache, with Win10CacheEntry as a typed
entry and a deliberate design note about why a cached None must
still count as a hit. It is tested. It has never been reused. Round 14
promotes it to caching.py, adds a
StaleWhileRevalidateCache subclass for the one call site that
needs it, and deletes six hand-rolled copies.
Design the end state first, then subtract toward it. When the work is done,
server.py contains this and nothing else:
| Part | Lines | What it holds |
|---|---|---|
| Module docstring + imports | ~120 | One import per owning module. No from x import name
re-exports — those are what round 12 abolished. 38 gone, 129
srv names to go. |
| The container | ~200 | Construct ~22 services with their real collaborators. One
def build_*_service() each, resolved per call, never
captured. |
| The port registry | ~80 | Bind the dozen facade objects the route ladders depend on. |
| Startup + background loops | ~80 | startup_tasks, startup_banners, and the
thread starts that take a sweep callable so the bundle is
rebuilt per iteration. |
| HTTP bootstrap | ~40 | Port, handler class, serve. Everything else is in
http_app/ already. |
| Blank lines, section comments, the design notes worth keeping | ~180 | This file has good comments. Keep the ones that explain a decision; drop the ones that narrate code that has left. |
| Total | ~700 | Against a 2,000-line ceiling. The gap is the slack. |
Anything you find yourself wanting to keep in server.py that is
not on that list belongs somewhere else. That is the entire test.
Seven patterns cover everything below. Name the one you are applying in the commit message — it is how the next shift knows whether you were moving code or changing shape.
| Pattern | Where it goes | What it kills |
|---|---|---|
| Facade | A dozen ports between http_app/ and the modules. |
169 free names reached through srv, and the re-export tax
on every future round. |
| Registry / Abstract Factory | SERVERS, SCANNERS,
RESTART_HANDLERS, HEALTH_CHECKS,
AGENT_CARDS become typed collections built by a
factory. |
~430 lines of literal, and the whole family of “entry N is missing key K” bugs. |
| Template Method | caching.py: TtlCache and
StaleWhileRevalidateCache. |
Six hand-written caches, 14 of the 20 module locks, 6 TTL constants. |
| Chain of Responsibility | Document resolution — _source_document_reference,
_source_document_path,
_resolve_local_supporting_document,
_usable_document_reference,
_slot_reference and friends are already a
try-this-then-that chain written as nested if. |
~610 lines and about 35 interlocking private helpers, replaced by an ordered list of small resolvers. |
| Command | RESTART_HANDLERS / RESTARTABLE_KEYS, and the
GET/POST route ladders themselves. |
A 23-line dispatch dict and two ~600-line if ladders. |
| Strategy | Statement preflight and document processing —
run_statement_preflight (157 lines) and
process_scanned_document (171) are both
one-function-per-document-kind wearing a switch. |
The two largest functions in the file. |
| Repository / Adapter | The intake record's persistence layer, and every direct
_rol_get_connection() reach from a render path. |
The reason notes has six stays for three functions:
rendering code talking straight to a DB. |
Fourteen rounds. The order is not by size; it is by what unblocks what. Rounds 12–14 remove almost nothing and make every later round cheaper. Take them in order.
The Left column is arithmetic on the round-11 baseline of 7,378. It
is not a prediction of wc -l, because a second agent ships
features into this file in parallel — round 12 removed 41 lines and
the file still ended 154 higher; round 13 removed 579 and the file fell 387.
Track the column as “refactor debt retired”, and report your
round's diff separately from the file size.
| # | Round | Pattern | Out | Left | Why here |
|---|---|---|---|---|---|
| 12 | Ports & the death of srv
landed 6c69a3f4 |
Facade | 41 | 7,337 | Enabling, and it worked: moving code now costs a port method, not a
permanent re-export. srv. 220→147 sites.
What it did. |
| 13 | Config out, typed landed d04b519d | Registry | 579 | 6,758 | Came in at 579 removed against ~530 predicted — the first round to beat its estimate, because round 12's ports meant the wiring went into a port method instead of back into the namespace. Found a live duplicate agent tile. What it did. |
| 14 | One cache class, six instances | Template Method | ~180 | 6,580 | Removes the least on the board and unblocks the most: the agent cluster is six caches deep and round 24 cannot start until they are one class. Round 14 in full. |
| 15 | Intake record persistence | Repository | ~480 | 6,100 | Round 11 already mapped this surface. _fold_event_into_intake
(122) is the largest single piece. |
| 16 | Intake & report HTML | Strategy | ~540 | 5,560 | Needs 15 first. build_recent_intake_html is 171 lines of
string building in a server module. |
| 17 | Statement preflight | Strategy | ~450 | 5,110 | run_statement_preflight is 157 lines; the doc-kind switch
is the seam. |
| 18 | Document processing | Strategy | ~390 | 4,720 | process_scanned_document 171 +
process_pdf_document 57 + reprocess_report
37, one strategy per kind. |
| 19 | Document resolution chain | Chain of Responsibility | ~580 | 4,140 | The densest knot in the file: ~35 private helpers, none over 70 lines, all calling each other. |
| 20 | Receipt lookup & report rows | Repository | ~570 | 3,570 | Absorbs the deferred open-receipt and
notes clusters, which were never clusters of their
own. |
| 21 | Expense entry, edit, resolve | Repository | ~620 | 2,950 | _fetch_expenses_by_ids is 118 lines and
submit_manual_receipt_entry 81. |
| 22 | Scanner hardware, trainer, dispatch | Registry + Facade | ~600 | 2,350 | Round 11 drew half this boundary already. Finishes the Mazda block. |
| 23 | Server management & restart registry | Command | ~700 | 1,650 | The ceiling falls here. SERVERS left in
round 13; this is the behaviour that read it. |
| 24 | Agent registry, health, model stats | Facade | ~650 | 1,000 | Only tractable after 14. Six caches, two locks, three payload builders. |
| 25 | Report status & the sweep | Strategy | ~300 | ~700 | _extract_report_attention_detail (86) and whatever is
still standing. Then delete the dead comments. |
Read those numbers as content, not as yield. Historically a
move returns 70–100% of its measured figure because wiring lands back
in server.py. That ratio is exactly what round 12 attacks: once
the routes talk to ports, the add-back goes into the container as one
registration line instead of into the module namespace as a permanent name.
If the burn-down is running behind by round 16, the diagnosis is almost
certainly that round 12 was done partially. It was not: the mechanism is
complete and the ceiling test will tell you the moment it starts leaking.
About 180 lines, and the smallest removal on the board. Take it anyway, and take it next: round 24 (the agent registry, health and model stats) is the second-biggest block left and it is pinned behind this one. Six hand-written read-through caches sit under it, each with its own dict, its own lock and its own TTL constant, and no round can restructure that cluster while the caching is re-typed inline at every call site.
Six copies of the same eleven lines. _agent_list,
_agent_activity, _agent_health,
_letta_id, _weekly_remaining and
_model_stats_agents each carry a
{'value': None, 'ts': 0.0} dict, a
threading.Lock() and a *_CACHE_TTL constant. That
is 14 of server.py's 20 module-level locks and 6 of its 7
hand-rolled TTL constants. Round 13 did not touch any of them — it
moved data, and these are state.
build_agent_list is the odd one and the interesting one. It
adds stale-while-revalidate on top: when the entry is stale it serves the
stale value immediately and kicks a background thread, because a cold
rebuild can block >10s on the Letta roster fetch and trip the browser's
fetch timeout. It stores the in-flight flag in the same untyped dict
as the value, so a rebuild that raises leaves refreshing stuck
True and the cache never refreshes again.
monitoring/win10_node.py contains _TtlCache with
Win10CacheEntry as a typed entry, and a deliberate design note
about why a cached None must still count as a hit
— the entry, not the value, is the presence sentinel. Five of the six
hand-rolled copies get that wrong: they test value is not None,
so any cache whose legal answers include None re-fetches every
single call while looking exactly like a working cache.
_TtlCache to caching.py
unchanged, with CacheEntry as the typed entry, and leave
monitoring/win10_node.py importing it. Do this as its own
step and keep tests/test_win10_node.py green before you
convert anything else — it is the only proof you have that the class
behaves, and it is proof written against the real caller._letta_id is the one to do first: it is the one whose
None is a legitimate cached answer (an agent absent from the
Letta roster) and it already carries LETTA_ROSTER_NEG_TTL and
a separate _letta_roster_fetched_at float precisely to work
around not being able to cache it. Check whether that whole negative-TTL
apparatus is still needed once the entry is the sentinel — if it is
not, say so, because it is a genuine simplification rather than a
relocation.StaleWhileRevalidateCache as a subclass
for build_agent_list, with the in-flight flag as a field on a
typed entry rather than a key in the value dict, and the flag cleared in a
finally:. That is the Template Method: the base decides
hit/miss/expiry, the subclass decides what a stale hit does.grep -n "'ts': 0.0\|time.time() -" server.py.None is a hit. One test per
converted cache, driving the underlying fetch through a counter, asserting
the fetch ran once across two calls that both answer None.
This is the defect the round removes and it is reachable in five of the
six today — show it as the old code, inline, per rule 10.build_agent_list fails this._agent_activity lock exists because an 11-agent sweep over the
DERP-relayed Letta API takes ~30s while the frontend polls every 5s. Drive
two threads at a cold cache and assert the fetch ran once. If that property
is not tested it will be lost in the conversion and nobody will notice
until the Agent Management tab starts taking 30 seconds.poll_loop).
The trap in this round. Round 13 hit rule 3's
second-binding failure twice, and round 14 is far more exposed to it because
it is moving state, not data. A cache object constructed in
server.py closes over whatever it was handed; a test that
monkeypatches server.get_letta_id afterwards will not be seen
unless the cache calls through the module at fetch time. The symptom is not
a red test — it is a green one that took 30 seconds and talked to the
live Letta API. Round 13's own notes on where this bit are in
the section below.
−579/+192 in server.py, which fell 7,533 → 7,146.
Module-level assignment went 704 → 262 lines and the ~430 lines
of config literal are gone. 3,162 tests pass, up from 3,047. It is the first
round to beat its own estimate (~530 predicted, 579 landed), and the reason
is round 12: the wiring left behind went into a port method instead of back
into the module namespace.
| Module | Models | Replaces |
|---|---|---|
finance/reporting_categories.py |
ReportingCategory |
the four parallel REPORTING_CATEGORY_* dicts |
hardware/scanners.py |
ScannerSpec | SCANNERS |
servers/registry.py |
ServerSpec, NamedCheckProbe,
HttpProbe, TcpProbe, LogOnlyProbe |
SERVERS |
servers/restart.py |
RestartCommand, RestartRegistry |
RESTART_HANDLERS + RESTARTABLE_KEYS |
agents/registry.py |
LettaAgentSpec, AgentCard,
VoiceOption |
LETTA_AGENTS, AGENT_CARDS,
AGENT_VOICE_OPTIONS |
finance/report_registry.py |
ReportMonth, FinanceReportSpec |
the two parallel month dicts, ROL_FINANCE_REPORTS |
intake/statuses.py |
TerminalIntakeStatus (Literal) |
_TERMINAL_INTAKE_STATUSES |
intake/progress.py |
MazdaProgressStep |
_MAZDA_PROGRESS_LABELS |
Every legacy name stayed as a derived view — the same dict, the same list, the same per-entry key order — so not one of the ~60 consumers and not one of the ~40 tests that monkeypatch them had to move. That is what made a 579-line removal near-zero risk.
LETTA_AGENTS listed
Shelia twice, identically. The roster is a list and
build_agent_list() iterates it, so /api/agents was
serving 21 tiles for 20 agents and Agent Management rendered two identical
Shelia cards. Confirmed against the live dashboard before the fix and again
after. AGENT_CARDS carried the same duplicate as a repeated
dict key, where Python silently keeps the last — which is why
the card text looked right and hid the roster bug entirely. This is number
seven of the round-6 postscript's family, and the
mildest: unlike the other six it was visible, and nobody had looked.
voice/config.py's
KNOWN_AGENT_NAMES says it is “kept in sync with
LETTA_AGENTS” and is not. Toyota — the
receptionist the voice path routes through — has no mishear correction
at all, and one minion is Suzuki Patch on the roster but
Suzuki Patcher in the voice list, so a spoken “Suzuki
Patcher” is corrected to a name no agent answers to. Editing
it changes what whisper hears, which is a behaviour change and does not
belong in a config-typing commit (rule 15). The exact drift is recorded in
agents/registry.py and asserted in
tests/test_agents_registry.py, so it cannot widen while it
waits. Whoever takes the voice pipeline next should take this.
Everything else in the round was a defence, and rule 11 says to say
so. The REPORTING_CATEGORY_* family guards a fallback path
superseded at runtime by ICategoryTaxonomy — exactly as
the diagnosis predicted. SERVERS, which the diagnosis suspected
twice, turned out to be clean: no mistyped unit, no dangling
depends_on, no unknown check name.
A model that only mirrors its dict is still worth writing — that is the Pydantic section's argument and it holds. But five of round 13's models check something the dict could never express, and those are the ones to copy the technique from:
ServerSpec.probe is a discriminated union.
The plan called this out in advance and it was right for a reason the plan
did not give: server_health() resolves check
before health_url, so an entry carrying both
advertised a health URL on /api/servers that was never pinged.
Note also where the plan's own prediction needed correcting — it
called the union “log_file / health_url / tcp_check / check”,
but the letta entry legitimately has both an HTTP
probe and a log. Log tailing is independent; the union is over the
active probe, and “no active probe” is its own case,
which is what makes log_file required exactly there.ScannerSpec cross-checks its own matchers.
namelike and driver_match must appear in
device. A near-miss matched the other HP on the same box and
turned the Diagnostics tab green for a scanner that cannot scan.ReportMonth requires the range to be the month the
key names. feb-2025 ending 2025-02-29 was
writable before, and its only symptom is a month-status query that answers
nothing.MazdaProgressStep pins position == step number.
_mazda_progress_from_messages indexes a parallel statuses list
by number (statuses[1], [2], [7]).
Those indices were correct only because the labels happened to be in order.RestartRegistry.check_covers. A Server
Management tile with no restart command renders without a Restart button,
which reads as a design choice rather than a missing registration. Note the
asymmetry it had to permit: chatgpt-provider keeps a handler
with its tile commented out, so the registry covers the tiles rather than
equalling them.
Rule 9 asks for the old literal and the new view side by side on real live
values. Round 13 did that and made it permanent: the tests read the
pre-refactor literals out of git at the baseline commit,
ast-evaluate them against this box's live values, and diff. A
git object is immutable, so the comparison cannot decay into a restatement of
the new code the way a pasted copy does — and it stays honest about the
values that differ per machine (PORT, the Letta URL, two log
paths). See tests/test_servers_registry.py.
Every view came back byte-identical including per-entry key order. The one
intended difference is the duplicate roster entry, and the live
/api/agents diff before and after the deploy shows exactly that
and nothing else.
Worth reading before round 14, which is more exposed to this than round 13 was:
_rol_finance_reports_for_month moved to the registry and
five tests went red immediately. ~15 tests drive the report paths by
monkeypatching server.ROL_FINANCE_REPORTS, and a function
closing over the registry's global cannot see that.
It was moved back — three lines in
server.py, deliberately not imported from the module that also
defines it, with a comment saying why. The registry keeps its own copy for
callers that want the real list.RESTART_HANDLERS became a derived view, so
monkeypatch.setitem on it began patching a copy. That test was
repointed at RESTART_REGISTRY, which is the
thing that dispatches. Opposite resolution to the first, same diagnosis:
ask which object owns the behaviour, then point the test at that.
DIRECT_SERVICES needed no entry this round, and
that is worth understanding rather than copying. Round 13 pointed routes at
ports, and every port adapter resolves through server at
call time — so ServiceRecorder's stubs still land. Point a
route at an owning module instead and they do not. Rule 18 stands.
147/116 → 134/109. Eight names left the
ladder: the seven config names the round moved
(SERVERS, RESTARTABLE_KEYS, LETTA_AGENTS,
ROL_FINANCES_REPORTS_MONTHS,
ROL_FINANCES_REPORTS_DEFAULT_MONTH,
ROL_FINANCES_REPORTS_URL_PREFIX,
_rol_finance_reports_for_month) plus get_letta_id,
which went with the receptionist lookup that was its only caller there.
The ceiling test was red on arrival, and that is the mechanism
working. A parallel feature (e96ace55) added
srv._resolve_expense_receipt_path without lowering anything, so
tests/test_http_app_ports.py was failing before round 13 wrote a
line. That name belongs to DocumentPort and leaves in round 20;
round 13 absorbed the rise into its own reduction. If you arrive to a
red ceiling, that is what it is for — find who added the name, note
which port owns it, and absorb it. Do not raise the constants.
Each replaces a question the ladder was answering by joining globals,
which is the same lesson ScannerPort.image_path() taught in
round 12. Ask what the caller is trying to do:
ReportsPort.resolve_month_key() — the ladder was
joining ROL_FINANCES_REPORTS_MONTHS and
ROL_FINANCES_REPORTS_DEFAULT_MONTH in two places to answer
“which month tab is this request looking at?”. Plus
cards_for_month() and a url_prefix property.ServersPort.all() and restartable_keys()
— round 23 populates the behaviour that reads them.AgentsPort.receptionist() — the one route reading
LETTA_AGENTS was scanning the roster for Toyota and then
resolving its id. One question, one method, and the hardcoded name moved
out of the route to sit beside the lookup.category_taxonomy_seed.py was left alone. It
is a fifth copy of the reporting-category data — the plan's
diagnosis said four. It was not folded into the new list on purpose: its
whole job is to be an independent restatement that
tests/test_category_taxonomy.py pins id by id, and deriving it
from the same source would make that equivalence test tautological. Its own
docstring says as much. Leave it.SCANNERS is still a dict at every call site.
The plan said “turn it into list[ScannerSpec] and exactly
one adapter line changes”. The adapter is right; the sixteen call
sites in server.py are not — they all do
cfg.get('output'), and ~10 tests monkeypatch
server.SCANNERS with plain dicts to drive unknown and
misconfigured scanners. The specs validate the literal at import, which is
where the mis-scan defect lives; converting the readers is round 22's, with
the hardware.server.py.
They are behaviour bound to that module's state. Round 23 moves them; round
13 only gave them a shape.js/. The month and scanner lists
are still hardcoded as default constructor arguments in
RolFinanceReportsController. The Python side is now one typed
collection, so serving it and having the JS read it is a clean follow-up
— and it is a separate commit (rule 15).srv name
landed 6c69a3f4
Kept because you will be extending this mechanism, and because one part of
it is a trap that stays invisible until it bites. Removed 41 lines from
server.py; the point was never the number.
| File | What it holds |
|---|---|
http_app/ports.py |
Fourteen Protocol classes, one per collaborator group.
Protocol, not ABC: the implementations are plain modules
and objects, checked structurally, so a later round can swap a
module-backed adapter for a real service class without touching a
route. ScannerPort is populated; the other thirteen carry
the names they will absorb in their docstrings and are filled in by the
round that moves their code. Resist a fifteenth —
a grab-bag MiscPort is srv renamed. |
http_app/registry.py |
One frozen Ports dataclass and current_ports(),
which resolves server out of sys.modules and
builds a fresh bundle per call. Ports carries a
field per populated port only; thirteen empty adapters built
per request would answer no question. Add your field when you populate
your port. |
tests/test_http_app_ports.py |
55 tests. Production wires the real object; the bundle is rebuilt per
call and a port constructed before a rebind still sees it; all 38
dropped names are asserted absent; and the srv. count is
pinned as a ceiling that only ever falls. |
from x import f in a route snapshots f and
neuters rule 3's patch target, so the ladders say
ssh_checks.cached_ssh_health(…). 38 of the 47 left
server.py entirely; 9 stayed because server.py
itself calls them (HERE, LETTA_BASE_URL,
REPO_ROOT, SSH_CONNECTIONS,
ValidationError, manual_entry,
model_stats, render_excel_for_browser,
claude_sdk_account_payload) — those imports do real
work, so they are not re-exports. The routes stopped reaching through
server for them all the same.ScannerPort end to end, the worked example:
eight names, one tab you can watch work. Two collapsed on the way.
SCANNERS and SCAN_TOOLS_DIR were reached by one
route, which joined them to answer one question — where is this
scanner's last image? That is image_path(), and the route
no longer knows a scanner spec is a dict. Ask what the caller is trying to
do before you give a port a method that just re-exposes a global.monkeypatch.setattr(server, 'pc_metrics', …) gets a
friendly AttributeError instead of a silent no-op.tests/http_app_harness.py makes route tests inert by replacing
every function on server. A route pointed at its owning
module is no longer covered by that. The symptom is not a red test —
it is a green one that took 30 seconds and SSHed to another box,
drove a scanner, or ran whisper.
DIRECT_SERVICES in that file maps each converted module to the
names it owns, and the stub is recorded under the bare name so
stubbed.called('pc_metrics') reads the same on either side of
the move. Every round that converts a port must add to it.
Sizes are how many of the original 167 srv names each port
absorbs. Eleven names get no port — they are plumbing, and
they went the same way as the 47.
| Port | Names | Owns · populated by |
|---|---|---|
ReportsPort | 24 | Report discovery, path aliasing, status classification, the month and receipt-only queries, and the three HTML builders. The biggest port, and the one most likely to want splitting once round 16 has moved the HTML out. Rounds 16, 20, 25. |
AgentsPort | 23 | Roster, cards, voices, models, OAuth accounts, activity, health, headless runs. Round 24 — only tractable after 14. |
ModelStatsPort | 16 | Stats sources, mute overlay, Codex sync, ChatGPT provider accounts. Round 24. |
ServersPort | 14 | The Server Management tab: status, logs, restart, deploy.
Round 23, after 13 moves SERVERS out. |
MonitoringPort | 12 | PC metrics, SSH roster, Win10 containers, failure classification.
Every name behind it already lives in monitoring/, so
round 12 pointed the routes straight at those modules — this port
may never need populating at all. |
DocumentPort | 11 | Receipt lookup, supporting documents, presence checks, Excel render. Rounds 19, 20. |
ExpensePort | 10 | Manual entry, stored-expense edit and search, notes. Round 21. |
IntakePort | 9 | The recent-intake record, the halt file, scanner intake lookups. Round 15. |
ScannerPort | 8 | Populated, round 12. Scanner hardware and its diagnostics tab — small, self-contained, and a tab you can watch work. |
PipelinePort | 8 | Document processing, statement break-up, Mazda fill, reprocess. Rounds 17, 18. |
VoiceNotesPort | 8 | Voice upload, speech synthesis, the receptionist strategy, note
commands. Round 12 pointed the routes at voice/ directly;
the port populates when the voice pipeline gets a composition
root. |
TerminalPort | 6 | PTY spawn and reap, WebSocket framing — already
terminal/pty_session.py and
http_app/websocket.py. Round 12 pointed
terminal_ws.py at both directly. |
CategoryPort | 5 | Recategorize, undo, vendor review. Already behind
finance/recategorize.py. |
MazdaPort | 2 | The mode switch. Round 11 reduced it to two verbs — it is the shape the other thirteen are aiming at. |
| no port | 11 | HERE, REPO_ROOT,
LETTA_BASE_URL, ValidationError,
_load_json/_append_json/_clear_json,
and the four Claude log files and locks. None of these is a service.
Import them from the module that owns them — and note that a route
holding a threading.Lock reached across a module boundary
is a bug waiting to be written, not an interface. Three of the
four Claude log locks are still reached through srv;
they are the ugliest thing left on this list. |
The justification is mechanical, not stylistic.
contracts.StrictModel is strict=True,
extra='forbid', frozen=True. A model that merely
mirrors the dict it replaces is therefore not decoration: it turns
all three of the silent failures this plan keeps finding — a coerced
value, an unexpected key, a mutation in flight — into exceptions at
the boundary. Mirroring is the safety. That is what makes writing
thirty of them defensible.
Against 152 return { boundaries and about 430 lines of config
literal, the dashboard currently has 35 StrictModel subclasses.
It also has 33 plain BaseModel subclasses
— a second, non-strict, non-frozen, extra-permitting base sitting
alongside the strict one. Audit those as you pass them; most should derive
from StrictModel, and the ones that genuinely cannot should say
why in a docstring.
All thirteen shipped, in seven modules. What follows is the worklist as it was written, with what each one actually turned out to be — the predictions were mostly right, and the two places they were wrong are the instructive ones.
| Model | Replaces | Predicted failure | What it turned out to be |
|---|---|---|---|
ReportingCategory |
4 parallel REPORTING_CATEGORY_* dicts (58 lines) |
A category in one dict and missing from another — a report row with no colour, served as if styled. | Defence, as predicted: the maps guard a fallback path
superseded by ICategoryTaxonomy. Also caught what the plan
missed — the ancestor map is not derivable from
db_id alone, so the model carries
ancestor_ids and validates that a bucket's own id rolls up
to itself. |
ServerSpec + four probes |
SERVERS (159) |
A mistyped systemd unit or host: a Restart button that reports success and does nothing. | Clean. No mistyped unit, no dangling
depends_on, no unknown check name. The union is over the
active probe, not over log_file —
see why. |
AgentCard | AGENT_CARDS (136) |
A card missing a key renders as a blank tile rather than an error. | Clean as copy, but it carried a repeated dict key
(Shelia) that Python resolved silently — which is what
hid the live roster bug next door. |
ScannerSpec | SCANNERS (23) |
A wrong device or namelike scans the wrong
hardware, or nothing, quietly. |
Defence, and the model went further than typing: it cross-checks that the matchers appear in the device string. |
RestartCommand + RestartRegistry |
RESTART_HANDLERS + RESTARTABLE_KEYS (24) |
A key in one and not the other. | Those two could never disagree. The real gap was that neither was tied
to SERVERS — so the registry checks
coverage of the tiles instead. |
LettaAgentSpec | LETTA_AGENTS (22) |
A malformed roster entry becomes
unknown-<name> downstream instead of raising. |
A live fix, and the round's headline. Not the predicted failure at all — a duplicated entry serving a duplicate agent tile. The finding. |
VoiceOption | AGENT_VOICE_OPTIONS (18) |
An unknown voice id reaches edge-tts and fails at speech time, on a background thread. | Defence. All sixteen well-formed. Led to the voice-config drift finding next door. |
FinanceReportSpec | ROL_FINANCE_REPORTS (16) |
A dir that does not exist becomes an empty tab. | Defence. The model checks the dir is a bare folder name, since it is joined to the month's base dir. |
ReportMonth |
ROL_FINANCES_REPORTS_MONTHS +
ROL_FINANCES_MONTH_RANGES (12) |
Two more parallel dicts on the same four keys. A month in one and not the other is a tab whose status query silently returns nothing. | Defence, plus the stronger invariant the plan did not ask for: the range must be the month the key names. |
MazdaProgressStep | _MAZDA_PROGRESS_LABELS (11) |
An unlabelled stage renders as a blank progress line. | Defence, and it found the sharper risk: the consumer indexes a parallel statuses list by position. |
TerminalIntakeStatus (Literal) |
_TERMINAL_INTAKE_STATUSES (10) |
Round 11 flagged this as the natural first Literal of the intake block. | Done. It is a vocabulary now, with the reason it matters written
down: an unknown status is dropped, leaving the document on
processing forever. |
Three of the worklist did not move, and the reasons are worth
keeping. HealthCheckSpec became
CheckName, a Literal in
servers/registry.py, because HEALTH_CHECKS maps
names to functions defined in server.py — the
vocabulary can leave, the bindings cannot until round 23. A test asserts the
two agree in both directions. ModelOption and
OAuthProviderAccount were left for round 24, which owns Model
Stats: AGENT_MODEL_OPTIONS is already consumed through
AgentModelOptionsService, so typing it away from that service
would model the list rather than the boundary.
The 33 plain BaseModel subclasses are still
unaudited. Round 13 added none — all thirteen of its models derive from
StrictModel, with exactly one documented relaxation
(RestartCommand sets arbitrary_types_allowed,
because a handler is a callable; it stays frozen and extra-forbidding). The
dashboard now has 48 StrictModel subclasses, up from 35.
Take these as each round reaches them; do not do a payload-typing round of its own, because a model written away from the code that builds it just mirrors today's bug.
| Round | Models |
|---|---|
| 15–16 | RecentIntakeRow, IntakeHaltRecord,
ScannerIntakeRef, MazdaProgress |
| 17–18 | StatementPreflightPayload, StatementRecord,
PipelineResult, ProcessedDocument |
| 19–20 | DocumentReference, SupportingDocumentView,
ReceiptLookupResult, ReceiptOnlyRow,
MonthStatus, RecentScanRow |
| 21 | ManualEntryResult, ExpenseEditResult,
StoredExpenseEvent |
| 22–23 | ScannerDiagnostic, ScannerRuntimeStatus,
ServerHealthPayload, DeployResult |
| 24–25 | AgentActivity, AgentHealth,
WeeklyRemaining, ModelStatsAgentRow,
CodeStatus, ReportAttention |
The discipline that does not relax: keep the golden-payload test against the prior literals, error strings and key order included, and where you can, run old and new side by side against real live values. Rounds 9, 10 and 11 all did it and all three caught something.
js/ is 35,400 lines and it is not in bad shape
structurally — there are already 39 files in
js/abstract/, 43 in js/implementation/ and 76 test
files, and the live entry point js/dashboard-boot.js is 157
lines. The convention is established and it works. The problem is
granularity: the interfaces are too few and too big, and three
implementation files have grown into the same shape as
server.py.
server.py
violates it in Python. If an interface's consumers each use a different
third of it, it is three interfaces.
| File | Lines | Interface | Verdict |
|---|---|---|---|
implementation/rol-finance-reports-controller.js |
1,300 | none | The worst file in js/. One class, 20 constructor
parameters, 9 endpoints, and at least eight responsibilities: month
tabs, month detail, recent report, scanner tabs, iframe loading, a
picker dialog, a headless Mazda UI, and polling. |
implementation/manual-entry-form.js |
1,365 | manual-entry.interface.js (470) |
The interface is genuinely good — real typedefs, pure decision
logic, no DOM, mirrors finance/manual_entry.py's rules. It
is simply doing five jobs. Split the interface, and the form follows
it. |
implementation/detail-renderers.js |
1,464 | detail-renderer.interface.js (44) |
Best-shaped of the three: 15 exports behind a small, correct
interface. It is fifteen renderers in one file. One file each, plus a
DetailRendererRegistry, and the interface does not change
at all. |
dashboard-boot-baseline.js |
3,101 | — | Referenced only by dashboard-baseline.html. Dead weight
in every grep and every search result. Confirm nothing serves it, then
delete both — or, if it is deliberately frozen, say so in a header
comment so the next agent stops re-reading it. |
ReportFetcher · MonthNavigator ·
ReportTableView · CategoryPickerDialog
· ReprocessCommand
Take CategoryPickerDialog first —
_ensurePickerDialog through _closePicker is
~95 self-contained lines with its own DOM subtree, and it is the one
piece with an obvious test.
FieldValidation (already exists as its own file) ·
VendorResolution · AmountEntry ·
SaveRequestBuilder ·
StoredFindingReader
The typedefs already cluster this way —
ManualEntryValidation, ManualEntryPrefill,
StoredFindingRow are three separate concerns sharing one
file.
Mechanical, low-risk, and it is the one that makes the
js/tests/detail-renderers.test.js (615 lines) legible
again. Registry replaces whatever switch
currently picks a renderer.
The convention already exists —
js/tests/manual-entry.interface.test.js (616 lines) and
js/tests/interface-workspace.test.js. A new
*.interface.js without a matching
*.interface.test.js is unfinished.
RolFinanceReportsController's constructor hardcodes, as default
arguments, the four month keys (jan-2025…apr-2025),
the two scanner keys (window, freezer) and the
Mazda agent id. Python holds the same four months in two parallel
dicts (ROL_FINANCES_REPORTS_MONTHS and
ROL_FINANCES_MONTH_RANGES), the same two scanners in
SCANNERS, and the same agent id in
MAZDA_AGENT_ID. That is three copies of the month list
and two of the scanner list, across two languages, none of them
checked. Round 13 makes the Python side one typed collection; the
JS side should then take its list from an endpoint rather than a default
argument. Until it does, one plan rule applies unchanged: one destination,
one definition.
Rule 15 used to say: stay out of js/ unless the bug is there,
because a concurrent dashboard-boot.js refactor runs in
parallel. That refactor has landed — the live boot file is 157 lines
and the module tree under js/boot/ is real. The rule becomes:
touch js/ deliberately, one interface at a time, and
still run git status js/ first, because another agent
may be mid-edit. Do not mix a JS split and a Python extraction in the same
commit; they fail differently and you want to be able to revert one.
The residual-coupling script below has picked every round since round 8 and it is still the right tool for ordering work inside a round. It is no longer how rounds are chosen: the burn-down is fixed and ordered by dependency, not by weekly measurement. Use the script to decide what travels with a block once you have committed to that block, and to check that a block really is as self-contained as this plan claims.
The right question it answers: which names stay behind in
server.py after the move? A name referenced only by the
cluster travels with it and costs nothing.
import ast
from collections import defaultdict
src = open('server.py').read(); tree = ast.parse(src)
funcs = {n.name: n for n in tree.body
if isinstance(n, (ast.FunctionDef, ast.AsyncFunctionDef))}
top = set(funcs)
for n in tree.body:
if isinstance(n, ast.Assign):
for t in n.targets:
if isinstance(t, ast.Name): top.add(t.id)
refs = defaultdict(set) # name -> top-level funcs that reference it
for fname, node in funcs.items():
for sub in ast.walk(node):
if isinstance(sub, ast.Name) and isinstance(sub.ctx, ast.Load) \
and sub.id in top:
refs[sub.id].add(fname)
def deps(name):
n = funcs.get(name)
return set() if n is None else {
s.id for s in ast.walk(n)
if isinstance(s, ast.Name) and isinstance(s.ctx, ast.Load) and s.id in top}
names = {...} # the cluster you are considering
while True: # pull in what only this cluster uses
ext = set().union(*(deps(n) for n in names)) - names
travels = {x for x in ext if not (refs[x] - names)}
if not travels: break
names |= travels
stays = sorted(set().union(*(deps(n) for n in names)) - names)
lines = sum(funcs[n].end_lineno - funcs[n].lineno + 1 for n in names if n in funcs)
print(lines, 'lines |', len(stays), 'stay behind:', stays)
print('cluster members:', sorted(names))
http_app/. It parses
server.py and nothing else, so every srv.<name>
call in the route ladders is invisible to it. Seven such names in round 8,
nine in round 9, zero in round 10, two in round 11 — and round 11's
own table said zero, because that column is a grep from the round before.
Always re-grep. Since round 12 the suite catches
half of this: tests/test_http_app_ports.py fails if
the srv. count moves in either direction. It still cannot
catch a name you moved out from under a live srv. caller,
because tests/http_app_harness.py stubs the service layer and
those tests pass against a name that no longer exists._chatgpt_failover['last_note'] became
last_failover_note(); round 11's two pokes at
_MAZDA_MODE_SERVICE became mazda_mode_status()
and set_mazda_mode(). There are 26 private names left
in the 169. Ask what the caller is trying to do before you give it
a port method that just re-exposes the private thing.record_agent_send_error would have dragged three cache globals
along and broken a route. Round 11's
_watch_intake_for_problems would have dragged the whole
Trainer escalation service into a dispatch module — it became an
injected port instead. Round 8's “Document Vision health” 83
lines with zero stays was a mirage over an already-extracted registry.
Grep every name you are about to move, and ask which cluster it would
belong to if both were already extracted.Rules 1–11 were paid for in production incidents and are unchanged except where noted. Rules 12–17 are new for the sub-2,000 phase.
http_app/ before dropping any name,
even when a table says zero — that column is always a grep from the
previous round. The HTTP tests stub the service layer, so they will not
catch a missing name either.server. A re-export is a second binding; the moved
code closes over its own module global, so
monkeypatch.setattr(server, 'X', …) isolates nothing
while looking exactly like it does. This has now bitten eight rounds
running. The corollary: when you move code, the tests that break are the
honest ones — the dangerous ones are those that keep passing.poll_loop). A
wrapper whose body is one collaborator call is not a smell — it is
how a neighbouring cluster's boundary gets drawn without moving it (round
11's watch_intake).LETTA_BASE_URL,
LETTA_DOCKER_HOST, classify_failure) is imported,
not injected. So is a vocabulary: if the destination module already spells
out the values you are about to re-declare, annotate against its alias. As
of this round the rule crosses the wire too — see the JS finding
about the month and scanner lists._probe_codex_usage; round 11
dropped DispatchEvidence and
InvalidExecutionMode. Assert the dropped names are absent so
dead re-exports cannot creep back. Round 12 retired the bulk of this
rule, which was its point: 38 names left and
tests/test_http_app_ports.py asserts each one absent by name.
What survives is the judgement call — when your round's last caller
leaves, check whether server.py still calls it before you
keep the import.rol_finances object, add one test
that runs the path unpatched against the real one. (See
the postscript.)min_length=1 is not “not blank” when the consumer
.strip()s.classify_failure() (or whatever downstream reads it) and
assert the label, not the text. Also assert the strings that should keep
their label still do — scrubbing that goes too far turns a real rate
limit into a generic error.srv. reference. Round 12's value is destroyed by
round 15 adding two. If a route needs something new, it needs a port
method — or, if the thing already has an owning module, an import of
that module. Enforced since round 12:
SRV_SITE_CEILING / SRV_NAME_CEILING in
tests/test_http_app_ports.py fail the build if the count
rises, and fail it again if the count falls without someone lowering the
constants. Lower them; a ceiling nobody moves becomes a lie the next round
inherits.js/ deliberately, one interface at a time.
Supersedes the old "stay out of js/". The parallel
dashboard-boot.js refactor has landed. Still run
git status js/ first — another agent may be mid-edit
— and never mix a JS split and a Python extraction in one commit.server.py. Not a
cache, not a lock, not a registry dict. There are 20 locks and 6
hand-rolled caches to remove; adding a twenty-first is how the file grew in
the first place. State belongs to an object in the module that owns
it.tests/http_app_harness.py
makes route tests inert by replacing every function on server.
Point a route at an owning module and that stub no longer lands, and the
symptom is not a red test — it is a green one that took 30 seconds
and talked to another machine. Add the module to
DIRECT_SERVICES in the same commit. The wider rule: a safety
net is aimed at a specific patch target, so moving the target is a
change to the net, whether or not anything turns red.git status again before you commit.
Two agents write to this tree. Round 12 had two of its unfinished files
swept into an unrelated commit by a concurrent git add -A,
and wc -l server.py moved by ~170 lines mid-round for reasons
that had nothing to do with the refactor. Stage deliberately, and report
your round by its diff, not by the file size.
finance/vendor_lookup.py called
VendorCategoryLookup.guess_vendor_key, a method
rol_finances has never defined — and the test that shipped
in the same commit monkeypatched a fake lookup object defining that same
imaginary method. The suite stayed green while every production caller raised
AttributeError, and the route swallowed it into a 200 page
reading “Scanner Report build error”, so monitoring saw success
too. Fixed in 603d2d78.
Round 8's docker ps bug was the same animal in a third costume:
a failure whose observable signature is identical to success. Round 9's
classifiers were the fourth — a payload nobody could read, reported as
an account with headroom. Round 10's is the fifth and the worst: a parked
credential nobody could read, treated as a wildcard that matched every
account, and then written back. Round 11's is the sixth, the mildest and the
most instructive: a status record that failed to land, whose only signature
was a document sitting on processing forever — a scanned
receipt that looks, on the page, exactly like one still being worked on.
All six were invisible to the suite. All six needed someone to run the real thing and look at what came back.
Number seven arrived in round 13, and it was not either of the two
this postscript predicted. The mistyped-SERVERS-entry
and the missing-REPORTING_CATEGORY_CLASS-key were both
hypothetical; typing them found nothing wrong with either. What was actually
in the tree was a duplicated roster entry serving a duplicate agent
tile — and it breaks this family's pattern in the one way worth noting:
it was not invisible. It was on the screen, on the live
dashboard, for anyone who counted the cards. Six failures whose signature was
identical to success, and then one whose signature was plainly wrong that
nobody had looked at. Both are the same lesson from opposite ends: run the
real thing and read what came back.
The two predicted failures are unfound rather than disproven. The validators for both now exist, so if either is ever written it fails at import instead of in a browser.
916dd739 Extract the HTTP layer, hardware I/O and stats extractors out of server.py6e3ddc96 Type the Model Stats, PC Monitor and terminal seams with Pydantic9b82f01b Name the health-check contract, and move the two biggest checks out of server.py6d7606d9 Extract the health-poller cache and vendor-review dialog backend out of server.py1351c3cc Extract the Server Management clocks and log-file reading out of server.py5f7d2059 Extract the category write and its undo out of server.py9552ce0b Extract the Win10 box and the three agent tabs out of server.py5ae369aa Make the Win10 container probe survive the remote shell47778b3f Extract the SSH connection checks and the provider quota probes out of server.py5b1fc498 Give a refused usage payload a noun in its error text039ad71b Extract the ChatGPT provider auto-failover state machine out of server.py78c8b9fa Extract the Mazda dispatch fork and the execution mode out of server.py6c69a3f4 Facade: put ports between http_app and server.py (round 12)d04b519d Registry: config out of server.py, typed (round 13)
Follow-ups that removed no lines from server.py but which the
next shift inherits: dbecf4fc (plan update, round 7),
e28bd5fd (re-fetch the plan when its tab opens),
b25955d5 (correct the plan's own stale numbers),
b3eba0e1 (stop the plan recording a hash that is stale on
arrival).
| Module | What lives there |
|---|---|
model_stats/ | sources, windows, usage_history (leak detection), last_good, reader, assignments, extractors |
monitoring/pc_metrics.py | PcMonitor, PcMetric (derives alert from level) |
monitoring/server_lifecycle.py | starting window + down/stale clock, ServerStatus |
monitoring/log_files.py | tail/age/mtime-probe, server_log_rows |
monitoring/win10_node.py | the Win10 box's reachability, containers, dockerd and the two recovery buttons; Win10CacheEntry, _TtlCache — round 14 promotes _TtlCache to caching.py |
monitoring/ssh_checks.py | the tailnet roster, the ssh and tailscale probes, the debounced health cache, the per-connection log tail, and the Test button |
monitoring/provider_usage.py | the zero-token quota pipeline; CodexUsage, ClaudeUsage, UsagePayloadError, shape_detail, PROVIDER_USAGE_PROBES |
monitoring/chatgpt_failover.py | the account-swap state machine: cooldown gate, standby verdict, token heal, swap, 90s sweep; StandbyCredentials, last_failover_note |
intake/mazda_dispatch.py | the fork between Mazda's LLM turn and a human's inbox: the mode branch, the scan dispatch, the late-delivery probe, and the record written when it declines; IntakeOutcome, Collaborators, HUMAN_ONLY_MODE_STAGE_MESSAGE |
intake/mazda_mode.py | the Automatic/Semi-Automatic vocabulary, the operator's switch and its store, and the MAZDA_DECISION_MODE parser: MazdaMode, ExecutionModeConfig, resolve_execution_mode, MazdaModeService, MazdaModeState |
intake/ (rest) | dispatch_evidence.py, scan_message.py, trainer_contracts.py, trainer_escalation.py, trainer_notifier.py, trainer_recovery.py |
http_app/ | the whole HTTP layer: transport, the two route ladders, terminal WS, runtime, static files, and srv, now being retired. ports.py holds the fourteen Protocols the ladders may depend on; registry.py holds current_ports(), which resolves through server per call and is the only module besides services.py allowed to name it. |
terminal/pty_session.py | pty spawn, process-group reap |
letta_code/runner.py | headless letta-code invocation + prompt validation |
health/ | probe.py (ProbeResult), failures.py, frita.py, document_vision.py, poller.py |
agents/ | letta_gateway.py, urllib_letta_gateway.py, model_options.py, message_views.py |
finance/vendor_review.py | list_vendor_keys, list_pending_vendor_review, set_receipt_vendor, PendingVendorReviewRow |
finance/recategorize.py | the category write, its undo, the report-row repaint, the undo journal's composition root, ReportRowClass, RecategorizeRequest |
hosts.py, paths.py, letta_ids.py | shared constants that previously had 2–4 independent env-read/fallback copies. LETTA_DOCKER_HOST reached four in round 10. |
finance/reporting_categories.py |
round 13. ReportingCategory and the thirteen reporting
buckets; the four legacy REPORTING_CATEGORY_* dicts are
derived views of it. NOT the runtime source —
ICategoryTaxonomy is; these guard the fallback.
category_taxonomy_seed.py is a fifth copy and is
deliberately not derived from this. |
hardware/scanners.py |
round 13. ScannerSpec; cross-checks that
namelike/driver_match describe the
device the script drives. SCANNERS is a derived
dict view — the sixteen call sites still read dicts, and round 22
converts them with the hardware. |
servers/registry.py |
round 13. ServerSpec with a discriminated
ServerProbe union (NamedCheckProbe,
HttpProbe, TcpProbe, LogOnlyProbe),
CheckName as the check vocabulary, and
build_server_specs() — a factory, because four entries
interpolate values the composition root owns. |
servers/restart.py |
round 13. RestartCommand + RestartRegistry
(Command). Owns dispatch and the check that every Server Management tile
has a Restart button. The handlers themselves stay in
server.py until round 23. |
agents/registry.py |
round 13. LettaAgentSpec, AgentCard,
VoiceOption, and the roster/cards/voice catalogue. Refuses a
roster with a repeated name or Letta id — which is how the live
duplicate-Shelia tile was found. Also records the
voice/config.py drift. |
finance/report_registry.py |
round 13. ReportMonth (the two parallel month dicts, with
the range validated against the key) and FinanceReportSpec.
MONTH_KEYS is what the JS controller should read instead of
its hardcoded default arguments. |
intake/statuses.py, intake/progress.py |
round 13. TerminalIntakeStatus (the intake block's first
Literal, flagged by round 11) and MazdaProgressStep, which
pins position == step number because the consumer indexes by
position. |
contracts.py | StrictModel — the one strict, frozen, extra-forbidding base every boundary model derives from. 48 subclasses after round 13, up from 35. 33 plain BaseModel subclasses still bypass it; audit them as you pass. |
ssh_gateway.py | dead — a second, subtly wrong SSH probe nothing imports. ConfiguredIdentityStrategy returns the first identity file that exists, which is the bug ssh_test's docstring exists to describe: a key can exist and no longer be authorized. Left in place (another agent's parallel work), but do not build on it, and port the fallback chain in first if a future round wants the adapter shape. tests/test_ssh_checks.py names the trap. |
ModelStatSource.kind: Literal['codex','claude','gemini'] — a typo used to render a card with no windows and no error.PcMonitor.memory_source: Literal['windows','linux'], frozen=True for the process-wide registry.PcMetric derives alert from level as a property, so a new level can't be added without deciding whether it blinks the tab.ProbeResult.hard — suppresses the Restart button for failures a restart can't fix; extra='forbid' catches misspelled flags.HealthCacheEntry — the poller reassigns a whole entry rather than mutating fields, so a corrupt debounce record can't be read back into a later poll cycle.PendingVendorReviewRow — pins the "pick a vendor" dialog's row shape so a renamed DB column fails at model construction, not as a blank cell.ServerStatus — the four-state tab vocabulary, validated where a fifth state would silently escalate a healthy server as stale.ReportRowClass — the only shape a cat-* class may have, checked at the one place that writes it into a report file on disk.RecategorizeRequest — a row id must be a whole number and an amount must be finite, because int(3.9) and Decimal('nan') both used to succeed and then quietly do the wrong thing.Win10CacheEntry — the entry, not the value, is the presence sentinel, so a cache whose legal answers include None can actually cache one.ConversationMessages — the one message path that bypasses ILettaGateway, where an unrecognised payload used to become an empty tab indistinguishable from a quiet agent.CodexUsage / ClaudeUsage — the payload that decides whether the fleet's account gets swapped. An unreadable body used to classify as headroom; the claude side rendered it as a confident 5h 0% / weekly 0%. Now refused, and the refusal deliberately worded so the dashboard does not call it a rate limit.StandbyCredentials — the parked second ChatGPT account. A bundle without account_id did not fail the "never park another account's token" guard, it disabled it, and the heal then overwrote the fleet's only spare credential with a stranger's single-use refresh token.IntakeOutcome — the record written for a document Mazda was never given. Three fields required because each one's absence silently switches something off: status_source, conversation_id, and dispatched_at > 0. The (status, status_source) pairing is validated too.MazdaMode (not a model) — one Literal alias replacing four independent spellings of the same two words. A vocabulary is worth de-duplicating for the same reason a hostname is.
Round 13's thirteen, in one line each.
ReportingCategory — a bucket's own id must roll up to
itself, so the picker cannot write a category the report then attributes
elsewhere. ScannerSpec — the diagnostics matchers must
appear in the device string the scan script drives. ServerSpec
+ its four probes — exactly one active probe, because
server_health() resolves check before
health_url and an entry with both advertised a URL it never
pinged. RestartCommand — a key and its handler stop being
two facts, and the registry must cover every tile.
LettaAgentSpec — no repeated name or Letta id, which is
what caught the live duplicate tile. AgentCard — a card
cannot be half-filled. VoiceOption — a voice id is checked
where it is chosen, not on a background thread at speech time.
ReportMonth — the calendar range must be the month the key
names. FinanceReportSpec — a report dir is a bare folder
name, not a path. MazdaProgressStep — position equals step
number, because the consumer indexes by position.
TerminalIntakeStatus — one Literal for the vocabulary
whose misses are silently dropped.
Round 12 added no model, and that is the honest report
(rule 11). It changed shape: fourteen Protocols and a frozen
Ports bundle. The one thing it typed by accident is worth
noting anyway — ScannerPort.image_path() replaced a route
that reached for srv.SCANNERS[key]['output'], so the browser
side of the scanner config now has exactly one reader. That is what makes
ScannerSpec cheap in round 13.
Next, per the worklist: round 14
adds no config model at all — it replaces six hand-rolled caches with
one class, and its typed entry (CacheEntry, promoted from
Win10CacheEntry) is the model that matters, for exactly the
reason Win10CacheEntry mattered: the entry, not the
value, is the presence sentinel, so a cache whose legal answers
include None can actually cache one. Five of the six get that
wrong today.
It is already wired in: dashboard.html has a
#dashboard-refactor-plan-frame iframe pointing at
/notes_plans_handoffs/dashboard_refactor_plan.html, behind the
Dashboard Refactor tab under Project Plans. Live at
https://desktop-2obsqmc.tailb8fc54.ts.net/.
A trap that looks exactly like a failed deploy: a tab switch only toggles a
CSS class, so the iframe is fetched once, when the dashboard page itself
loads. A dashboard left open all day keeps showing whichever revision it
pulled at load time, however many times you redeploy. The frame carries
data-refresh-on-show, and activateView in
js/boot/view-navigator.js re-fetches marked frames every time
their tab is opened. If a plan still looks stale, reload the dashboard page
before suspecting the deploy — and confirm with
curl -s <live>/notes_plans_handoffs/<plan>.html | grep <marker>,
which asks the server rather than the browser.