First Edition · Revised through the Nine Factory Families · Internal circulation
Front Matter
This manual replaces the change-log that previously occupied this page. A change-log answers the question "what happened last?"; a manual answers the more durable question "how is this thing built, and how do I work on it?" The project has matured past the point where a running diary serves it. What follows is organized the way a course text is organized: foundations first, architecture second, daily practice third, reference material last.
The subject is Mazda — a Letta agent that verifies finance data (bank statements and scanned receipts) inside the larger rol_finances project — together with the framework that surrounds her. That framework has a single, unfashionable thesis, stated here once and defended throughout:
We improve the wrapper around a cheap large language model, and we never touch the model itself. The wrapper is everything we can version and roll back: system messages, prompt templates, tool descriptions, context strategy, memory notes, and workflows. The model is a fixed, interchangeable engine. All of our engineering effort — tracing, judging, proposing, A/B testing, gating, activating, rolling back — operates on the wrapper.
A reader who internalizes only that paragraph has the core of the design. The rest of the manual explains the machinery that makes the thesis operational and safe.
Part I establishes what Mazda is and the live system she runs in. Part II is the heart of the book: the three-layer architecture, the nine factory families that compose a runtime, and the self-improvement loop that closes around them. Part III is operational — how to build, test, and run, and how to verify the live fleet. The appendices are reference catalogs.
Each chapter ends with Exercises. They are not busywork: each one corresponds to a real task a developer on this project will eventually perform. Working them is the fastest way to become productive.
— The maintainers, Autonomous Systems group
Front Matter
In which we fix the subject — what Mazda is, what she is not, and the live system in which she runs — before any line of architecture is drawn.
Chapter One
Mazda is an orchestrator, not a monolith. She is a single Letta agent who delegates the real work to a small team of specialist agents, each of which drives a Claude Agent SDK session for one slice of the finance-verification problem. Understanding this division of labor is the prerequisite for everything else.
It is easy to conflate two things that share the name "Mazda." Keep them distinct:
agent_self_improvement) that runs an agent, judges it, proposes wrapper changes, A/B tests them, and activates winners with rollback. This is the control plane. Parts II–III describe it. It is finance-blind and agent-blind: it reaches Mazda only through ports.The wrapper is the versioned, roll-back-able envelope around a fixed LLM: its system messages, prompt templates, tool descriptions, context-selection strategy, memory notes, and workflow definitions. The framework improves the wrapper; it never fine-tunes or swaps the model as an act of "improvement."
Mazda is the orchestrator — there is no separate "Orchestrator" agent. She delegates to five minions via Letta's inter-agent messaging. Each minion's specialty is driving a Claude Agent SDK (TypeScript) session for its slice of work. Finance domain logic — totals, vendor key, category, and duplicate verifiers — is exposed to those SDK sessions as MCP tools, not called by Mazda directly.
run_claude_code_sdk tool fans out to a single executor; the finance verifiers reach the SDK session over MCP.Two earlier designs were considered and deliberately discarded. A new developer will find references to them in old commits and should not revive them:
(a) The "give Mazda seven direct finance tools" design, in which the orchestrator called verifiers herself. Superseded by the minion + MCP arrangement above. (b) The separate-TypeScript-package minion path (IClaudeAgentSdkRunner / DefaultMinion / cli.ts) and its Python adapter. The live run_claude_code_sdk tool (one Python tool per agent, inline TS through the :8799 executor) is the canonical Phase 2 implementation; the package path is retired dead code.
The adjective is earned by the control plane, not by any single clever prompt. The framework can observe Mazda doing her job, judge whether she actually succeeded (not merely whether she produced output), propose a concrete edit to her wrapper, run that edit against a baseline behind quality gates, and either activate it or discard it — all with a snapshot it can roll back to. Chapter 5 develops this loop in full. The thesis of the Preface is what keeps it honest: every "improvement" is a wrapper revision, never a model change.
Chapter Two
A framework is only as real as the infrastructure it runs against. This chapter is the field guide to the live deployment: the agents and their identifiers, the executor that backs their SDK sessions, and the finance MCP server they call. Treat the identifiers here as authoritative reference, not as prose to be read once.
Mazda is the Letta agent agent-6b536cf4-ec88-4290-b595-fed21d14bd8e. She renders six system/ memory blocks into compiled context: persona, human, db_schema, environment, team_agents, and verification_procedure. (How to verify that rendering — the one authoritative method — is §7.2.)
All five are git-memory-enabled, each renders system/persona + system/human, and all five share a single run_claude_code_sdk tool — one tool identifier reused across the team, so one client.tools.update covers them all.
| Minion | Agent ID | Specialty |
|---|---|---|
mazda-router-agent | agent-bc561f63-a5bd-4192-806e-58d92593da2b | Task routing |
mazda-parser-agent | agent-a5063757-46c7-4054-a07d-2b1263db43a8 | Parsing / structured extraction |
mazda-vendor-identity-agent | agent-acd624ac-17f2-4a74-aa34-78036cac4d66 | Vendor normalization |
mazda-receipt-linker-agent | agent-9a14f800-d848-4914-bfd4-53ab62bc177b | Receipt ↔ transaction linking |
mazda-categorization-agent | agent-c429ff25-c8af-4f1a-a6f1-6d48307e2874 | Category resolution |
Each minion's run_claude_code_sdk(task, context, working_dir) writes an inline TypeScript file that calls claude().withModel('sonnet').allowTools('Read','Write','Edit','Bash','Glob','Grep').inDirectory(workDir).query(prompt).asText() and runs it through the executor. The minions are built by setup_mazda_minions.py on the Windows host.
The executor is the small HTTP service that actually spawns a Claude SDK session. It is published at http://100.80.49.10:8799/claude_sdk (fallback 172.17.0.1:8799). It authenticates with a bearer token carried in the tool source; an unauthenticated POST returns 401, an empty body returns 422, and a bare GET returns 405 — the three responses that together prove "alive and validating."
A round-trip with the task "reply PONG" returns {"status":"ok","output":"PONG"}. This is the cheap smoke test that confirms a real Claude SDK session ran end-to-end, not just that the request schema was accepted.
Finance domain logic reaches the SDK sessions through a FastMCP("finance_verifiers_mcp") server that exposes four deterministic verifier tools over stdio. It is wired into the executor's inline-TS template via .withMCP(...) — covering all five minions at once — and is deployed as a systemd service via deployment/mazda-tools-mcp.service.
| Tool | What it checks | Status |
|---|---|---|
verify_statement_totals | Line items sum to the stated total | live |
check_vendor_key | Vendor key is recognized (vendor_category.yaml) | live |
check_category | Category name is valid | live |
check_duplicates | Row is not already in the finance DB | needs DB creds |
check_duplicates requires pymysql plus finance MySQL credentials; absent those on a given host it returns a structured {"error": …} rather than crashing. The other three are pure Python and always available.
client.tools.update calls does it take to update run_claude_code_sdk across all five minions, and why?The three layers, the nine factory families that compose a runtime from them, and the self-improvement loop that closes around the whole.
Chapter Three
The framework is built in three layers, strictly ordered by the direction in which dependencies are allowed to point. The rule is simple and it is enforced by tests: program against contracts; concrete code lives behind them. Violating the layering does not merely offend taste — it breaks the test suite's sys.path setup.
| Layer | Package | Contents | May import |
|---|---|---|---|
| Contracts | contracts/ |
Pydantic value objects + ABC ports. The public surface. | Only the standard library + Pydantic. Never finance or agent packages. |
| Implementations | implementations/ |
Concrete adapters, grouped by milestone: persistence, runtime, evaluation, improvement, MCP, reporting, stubs. | Contracts; its own siblings. Finance code only by injection. |
| Agent packages | agent_packages/ |
Per-agent glue — Mazda's task ports, models, and the multi-agent finance orchestrator. | Contracts + implementations. This is where agent-specific detail is allowed to live. |
A port is an abstract base class in contracts/ describing a capability (e.g. ITraceRepository). An adapter is a concrete class in implementations/ that satisfies a port (e.g. SqliteTraceRepository). Callers depend on the port; the choice of adapter is made in exactly one place — the composition root of Chapter 4.
The kernel and contracts/ must never import finance packages or any agent package. Mazda is reached only through the generic IAgentPackageFactory and through finance objects injected as Any-typed constructor arguments. The Live*Verifier classes are the model to follow: they wrap real finance services by dependency injection (a VendorCategoryLookup, a DuplicateChecker), never by import.
This is why the evaluation layer ships two verifier families: in-memory stubs (KnownVendorKeyVerifier and friends) for unit tests that touch no infrastructure, and Live*Verifier adapters that wrap the real finance services by injection. The contracts never know which is in use.
The convention is invariant across the whole project: new behavior starts as a contract, then gets a concrete implementation behind it. A Pydantic model or an ABC goes into contracts/ first; the adapter follows in implementations/; commonly used names are re-exported from the relevant __init__.py. This discipline is what makes the factory families of the next chapter possible — you cannot manufacture what you have not first specified.
from nonprofit_finance import ... to a file in contracts/. State which rule it breaks and what mechanism will catch it.Chapter Four
This is the chapter the architecture is built around. The framework composes a runtime from families of related objects using the Abstract Factory pattern. Nine factories cover the nine subsystems; a tenth — the agent-package factory — joins them inside a single runtime family. A composition root wires the family from a profile, and an agent kernel runs tasks through it. By the end of this chapter you will be able to add a backend, swap a whole family, or trace one task from request to persisted verdict.
The kernel must create many related objects — a trace repository, a verdict judge, an event bus — without binding to their concrete classes. The Abstract Factory pattern lets us swap a whole compatible family (say, SQLite persistence + a cheap-LLM client + finance evaluation) without the kernel ever changing. The cost is one extra layer of indirection; the benefit is that "which backend" is a decision made in exactly one place.
Each factory is an ABC in contracts/factories.py; each create_* method returns a port, never a concrete type. The full method-by-method catalog is Appendix A; the families themselves are Table 4.1.
| # | Factory (ABC) | Manufactures | Concrete implementation |
|---|---|---|---|
| 1 | IPersistenceFactory | schema manager, unit of work, 5 repositories | SqlitePersistenceFactory(db_path) real |
| 2 | ILlmFactory | client, request builder, parser, token counter, budget guard | StubLlmFactory(model_name) stub |
| 3 | IPromptFactory | system-message provider/builder, template provider, composer, validator, diff | StubPromptFactory stub |
| 4 | IToolFactory | registry, adapter, permission policy, execution logger | StubToolFactory stub |
| 5 | IContextFactory | selector, budgeter, pack builder, memory/note readers + writers, staleness detector | StubContextFactory stub |
| 6 | IWorkflowFactory | workflow runner | StubWorkflowFactory stub |
| 7 | IEvaluationFactory | evaluator, verdict judge, classifier, regression detector, gate chain, 4 finance verifiers | FinanceEvaluationFactory(...) real |
| 8 | IImprovementFactory | proposal generator, experiment runner, approval gateway, snapshot store, activation + rollback | DefaultImprovementFactory(...) real |
| 9 | IReportingFactory | event bus, trace report builder, dashboard builder, regression report builder | DefaultReportingFactory real |
| + | IAgentPackageFactory | Mazda / Scissari / Frita packages by name | DefaultAgentPackageFactory real |
Not every subsystem has a production backend yet, and the architecture does not pretend otherwise. Four families — persistence, evaluation, improvement, reporting — have real implementations wired to SQLite and the finance verifiers. Five families — LLM, prompt, tool, context, workflow — currently ship stubs in implementations/stubs/.
A stub is a minimal but honest port implementation. Some stubs do real, trivial work (the workflow runner genuinely iterates steps in sequence; the request builder genuinely assembles a request); others raise NotImplementedError for operations that need infrastructure we have not built (an LLM complete(), a memory write()). A gap-filler is a stub that stands in for a port the real family hasn't implemented yet — e.g. InMemoryUnitOfWork and InMemoryEvaluationRepository, used by the SQLite persistence factory for its two not-yet-persisted ports.
This honesty matters at runtime: the kernel expects some ports to raise NotImplementedError and degrades gracefully (§4.6). Stubs are not technical debt to be hidden; they are placeholders with a precise contract, and they make the whole family composable today.
A single profile uses one factory from each family. The IAgentRuntimeFamily port bundles all ten as read-only properties, so a subsystem can be swapped by swapping one property's factory. The concrete bundle is DefaultAgentRuntimeFamily; it takes the ten factories as keyword arguments and exposes them.
class IAgentRuntimeFamily(ABC):
"""A bundle of compatible factories for one runtime profile."""
@property
@abstractmethod
def persistence_factory(self) -> IPersistenceFactory: ...
@property
@abstractmethod
def llm_factory(self) -> ILlmFactory: ...
# ... prompt, tool, context, workflow, evaluation,
# improvement, reporting, agent_package ...
If the kernel must never name a concrete class, something must. That something is the composition root. It is the single seam where abstraction is traded for concreteness, and it is deliberately the only such seam in the system.
DefaultCompositionRoot.build_kernel(profile) reads a RuntimeProfile and instantiates every concrete factory, bundles them into a DefaultAgentRuntimeFamily, and returns a DefaultAgentKernel. It is the only file in the framework permitted to import concrete factory classes.
The wiring order is not arbitrary — it encodes the cross-family dependencies. Persistence is built first (and its schema created) because the improvement factory needs persistence's wrapper-revision and improvement repositories injected into it. Stubs come next, evaluation is configured from profile.extra, then improvement is cross-wired, and reporting last.
class DefaultCompositionRoot(ICompositionRoot):
def build_kernel(self, profile: RuntimeProfile) -> IAgentKernel:
db_path = profile.db_path or "agent_improvement.sqlite3"
persistence = SqlitePersistenceFactory(db_path=db_path)
persistence.create_schema_manager().create_schema() # schema first
llm = StubLlmFactory(model_name=profile.llm)
prompt = StubPromptFactory()
tool = StubToolFactory()
context = StubContextFactory()
workflow = StubWorkflowFactory()
evaluation = FinanceEvaluationFactory( # configured from the profile
known_vendor_keys=set(profile.extra.get("known_vendor_keys", [])),
vendor_category_map=profile.extra.get("vendor_category_map", {}),
...)
agent_package = DefaultAgentPackageFactory()
improvement = DefaultImprovementFactory( # cross-wired with persistence
db_path=db_path,
wrapper_repo=persistence.create_wrapper_revision_repository(),
improvement_repo=persistence.create_improvement_repository())
reporting = DefaultReportingFactory()
family = DefaultAgentRuntimeFamily(
persistence=persistence, llm=llm, prompt=prompt, tool=tool,
context=context, workflow=workflow, evaluation=evaluation,
improvement=improvement, reporting=reporting,
agent_package=agent_package)
return DefaultAgentKernel(family=family, profile=profile)
The RuntimeProfile is a pure data record — it names the families ("sqlite", "gemini_flash_lite", "mazda", "finance_rules") and carries an extra dictionary for evaluation configuration and a db_path. The canonical first profile is sqlite_gemini_flash_lite_mazda_default. Selecting a different backend is, by design, an edit to a profile — not to the kernel.
The kernel depends only on IAgentRuntimeFamily. It never learns whether persistence is SQLite, whether the LLM is one model or another, or who Mazda is. Its run_task is short enough to read in full, and reading it is the best way to see the families cooperate.
DefaultAgentKernel.run_task. Four of the ten families participate; the verdict and event steps degrade gracefully when a stub raises NotImplementedError.The kernel wraps the verdict and event-publishing steps in try / except NotImplementedError. This is what lets a runtime composed of real persistence but stubbed reporting still produce a valid RunOutcome. The trace is always saved; the verdict is saved when an evaluation backend exists. The composition stays whole even while half its families are stubs.
A crucial property for anyone maintaining existing code: the factory layer was added on top of the direct-construction code, not in place of it. Every entry point and every pre-existing test that wires objects by hand continues to work unchanged. The factories are a second, optional path to the same concrete classes — proven by a dedicated test that constructs objects both ways and asserts they are equivalent.
Suppose the evidence store must move from SQLite to MySQL. The Abstract Factory pattern makes the blast radius precise:
MySqlTraceRepository and siblings as adapters behind the existing persistence ports.MySqlPersistenceFactory(IPersistenceFactory) returning them.profile.persistence == "mysql" to choose the factory.The kernel, the runtime family, every evaluator, and every test that depends on the persistence port are untouched. That containment is the entire return on the indirection the pattern costs.
RunRequest through run_task and list, in order, every factory whose create_* method is called. (Answer: see Figure 4.1.)GeminiLlmFactory that replaces the stub. Which existing files change, and which do not?Chapter Five
With a runtime composed and a kernel able to run one task, we can close the loop that gives the project its name. The loop runs the agent, judges whether it truly succeeded, proposes a wrapper edit, A/B tests that edit behind gates, and activates the winner with a snapshot it can roll back to. It imports no finance code and names no agent; it reaches the agent only through ports.
The kernel runs a task and persists a TraceRecord: inputs, outputs, tool calls, token usage, and the four wrapper revisions in force. This is the raw evidence everything downstream reasons over.
Judging answers "did the agent actually succeed?" — not merely "did it return text?" It is deterministic-rules-first: the finance verifiers produce structured evidence, a rule-based evaluator scores it into a ScoreCard, and the verdict judge assigns pass / fail / needs-review. An LLM judge strategy is consulted only for genuinely ambiguous cases (an unrecognized vendor, a fuzzy duplicate). A non-passing verdict is then labeled with a finance FailureType.
Cheap, deterministic rules decide every case they can. The expensive, non-deterministic LLM judge is a fallback for the residue of ambiguity — never the first resort. This keeps verdicts reproducible and keeps cost down.
From failure patterns, a proposal generator suggests a concrete wrapper edit. Edits are modeled as commands: PromptPatchCommand, ToolDescriptionPatchCommand, MemoryNoteCommand, ContextRulePatchCommand. A command is a reified, inspectable, reversible change — exactly the granularity the gates and rollback need.
A proposed edit is not trusted on its say-so. A baseline-versus-candidate experiment runs both wrappers and a regression detector compares their scorecards, so an "improvement" that quietly regresses some other case is caught before it ships.
The candidate must clear a GateChain of four independent gates before it is eligible for activation:
| Gate | Question it asks |
|---|---|
SafetyGate | Does the change violate a safety invariant? |
CostGate | Does it blow the token / latency budget? |
RegressionGate | Did any previously-passing case regress? |
UsefulnessGate | Is the improvement actually material? |
A gated winner is activated through a human approval gateway, with a snapshot taken first. If the change misbehaves in practice, the rollback service restores the known-good wrapper revision by id. Approval and activation records land in the same evidence store as everything else, so the whole history is auditable.
The loop is finance-blind infrastructure; only the thing being wrapped is Mazda-specific. The mapping is direct:
| Loop concept | Mazda usage |
|---|---|
IAgentRunner / run_task | Run Mazda's orchestration step; collect each minion's SDK-session result. |
| Trace / evidence store | Record Mazda→minion delegations and each SDK session (inputs, outputs, tool calls, tokens). |
IEvaluator family | Orchestration verdict; the finance verifiers become one pluggable evaluator. |
WrapperRevision | Versioned delegation logic + per-minion SDK task templates (prompts, allowed tools, params). |
| Proposal / Experiment / Gate / Rollback | Propose, A/B, gate, and human-approve a change to delegation strategy or a minion template — rollback-safe. |
RegressionGate. What happened, and what does the loop do next?WrapperRevision in this project.From architecture to the keyboard — how to build, test, and run the framework, and how to verify the live fleet without re-diagnosing solved problems.
Chapter Six
This chapter is the lab manual. It assumes you are working in tools/self_improving_agent/ and that the virtual environment lives at the repository root in ../../.venv, not inside this subdirectory — a detail that trips newcomers and is worth memorizing.
Use .venv/bin/python explicitly; the system python may be absent and python3 may lack pytest. The conftest.py + pytest.ini pair makes the package importable regardless of which directory you invoke from.
# everything (live integration tests self-skip if env unset)
../../.venv/bin/python -m pytest
# one file — e.g. the factory-family suite
../../.venv/bin/python -m pytest agent_self_improvement/tests/test_phase_02b_factory_families.py
# one test by name substring
../../.venv/bin/python -m pytest -k "vendor_key"
# unit only / certification-blocking integration only
../../.venv/bin/python -m pytest -m "not integration"
../../.venv/bin/python -m pytest -m p0
Unit tests (tests/test_phase_NN_*.py) use duck-typed fakes — no MySQL, no Letta, no network — and should always pass. Integration tests (tests/integration/p0|p1|p2) are marked and skip themselves unless the right environment variables are set. The severity tiers are p0 (blocks certification), p1 (blocks multi-agent expansion), and p2 (hardening).
The phase-numbered test files double as a curriculum: each maps to a milestone of the architecture. The factory work of Chapter 4 lives in its own phase file.
| Test file | Covers | Chapter |
|---|---|---|
test_phase_00_value_objects.py | Pydantic value objects, typed ids | 3 |
test_phase_01_database_contracts.py | SQLite evidence store | 3, 5 |
test_phase_02_wrapper_bundle.py | Wrapper revisions + snapshot/rollback | 5 |
test_phase_02b_factory_families.py | The nine factories, runtime family, composition root, kernel (18 tests) | 4 |
test_phase_03_agent_runner.py | Agent runner port + fakes | 4 |
test_phase_04_verdicts.py | Verifiers → scorecard → verdict → failure type | 5 |
test_phase_05–07 | Proposals, experiments, activation + reports | 5 |
test_phase_08–09 | Approval audit, concurrency guard | 5 |
Across the suite there are over two hundred test functions; the factory-family file alone contributes eighteen, proving that every factory creates its ports, the runtime family bundles them, the composition root builds a kernel, the kernel runs a task with a persisted trace, and direct instantiation still works.
mazda_verify_receipt.pyRuns Mazda's deterministic verifiers over one parsed receipt JSON, optionally asks the live Letta agent for a narrative, persists a trace + verdict to SQLite, and writes an HTML report to receipt_runs/.
python3 mazda_verify_receipt.py -f walgreens_2025-01-22.json # full run
python3 mazda_verify_receipt.py -f receipt.json --no-letta # skip live Letta turn
python3 mazda_verify_receipt.py -f receipt.json --db /tmp/run.sqlite3 --output out.html
mazda_run_team.py../../.venv/bin/python mazda_run_team.py --check-ready # verify specialist IDs
../../.venv/bin/python mazda_run_team.py --dry-run --source-uri /tmp/test.pdf # offline workflow
../../.venv/bin/python mazda_run_team.py --live --source-uri /tmp/test.pdf # live specialists
implementations/finance_mcp/server.py exposes the four verifier tools over stdio; it is deployed as a systemd service via deployment/mazda-tools-mcp.service behind mcp-proxy on port 8791.
tools/self_improving_agent/, write the exact command to run only the factory-family suite. Why must you spell out the venv path?p0 tests, and why is that correct behavior rather than a failure?--no-letta. Which stations of the Chapter 5 loop still execute, and which are skipped?Chapter Seven
The live fleet has a small number of recurring operational questions and exactly one correct way to answer the most important of them. This chapter records those answers so that a new instance never re-diagnoses a problem the project has already solved.
http://100.80.49.10:8283. Prefer this; it works from anywhere.docker-visible letta-postgres container is a stale 4-agent stack; never query it for live data.http://100.80.49.10:8799/claude_sdk (fallback 172.17.0.1:8799).The single authoritative test of whether an agent renders its memory is the <projection> count from a recompile. Everything else is a red herring.
curl -s -X POST http://100.80.49.10:8283/v1/agents/<AGENT_ID>/recompile \
| grep -oE '<projection>[^<]*</projection>'
# Each line = one system/ block rendered into context.
# Mazda = 6 · each minion = 2 · zero = amnesiac.
The agents.system DB column is the base prompt template — block content is injected at compile time and is never stored there, so it will always look "empty" of memory. An empty memfs repo (zero commits) is likewise fine: on this deployment the system/-labeled block drives rendering. And /v1/agents/<id>/context's system_prompt field is stale for dormant agents. Trust only the recompile projection count.
A recompile reads freshly-attached blocks on the next call, not the same instant they attach. If the first projection grep returns empty right after attaching a block, re-run it once.
Editing executor_server.py in place changes its inode, which breaks docker-desktop's single-file bind mount and makes docker restart fail with "no such file or directory." Recovery is to recreate via the idempotent deploy_frita_executor.sh. Directory mounts do not have this problem.
IS_SANDBOXThe executor runs as root, so claude --dangerously-skip-permissions is refused unless IS_SANDBOX=1 is set. It is now exported both in the SDK subprocess environment and as a container variable, so .skipPermissions() works under root.
memory_system_plan.html — the memory policy: author memory as system/** memfs files; blocks are a read-only projection; never raw POST/PATCH /v1/blocks.agent_self_improvement_plan.html — the design source for the control-plane loop and the nine factory families. (Its older "Mazda gets seven finance tools" framing is superseded; see §1.3.)docker restart frita-executor fails with "no such file or directory." What did someone just do, and what is the recovery command?IS_SANDBOX=1? What would silently fail without it?Appendix A
A method-by-method reference for the ten factories of Chapter 4. Every method returns a port (an ABC), never a concrete class.
| Factory | create_* methods |
|---|---|
IPersistenceFactory | schema_manager · unit_of_work · trace_repository · wrapper_revision_repository · prompt_artifact_repository · tool_artifact_repository · evaluation_repository · improvement_repository |
ILlmFactory | client · request_builder · response_parser · token_counter · budget_guard |
IPromptFactory | system_message_provider · system_message_builder · template_provider · composer · validator · diff |
IToolFactory | registry · adapter · permission_policy · execution_logger |
IContextFactory | selector · budgeter · pack_builder · memory_reader · memory_writer · note_reader · note_writer · staleness_detector |
IWorkflowFactory | runner |
IEvaluationFactory | evaluator · verdict_judge · failure_classifier · regression_detector · gate_chain · totals_verifier · vendor_key_verifier · category_verifier · duplicate_verifier |
IImprovementFactory | proposal_generator · experiment_runner · approval_gateway · snapshot_store · activation_service · rollback_service |
IReportingFactory | event_bus · trace_report_builder · dashboard_builder · regression_report_builder |
IAgentPackageFactory | create_agent_package(name) · supported_agents() |
| Contract | Role |
|---|---|
RuntimeProfile | Data-only record naming each family + db_path + extra. |
IAgentRuntimeFamily | Bundles all ten factories as read-only properties. |
ICompositionRoot | build_kernel(profile) → IAgentKernel; only place that names concrete classes. |
IAgentKernel | run_task(request) → RunOutcome. |
RunOutcome | { trace, verdict? } — the product of one run. |
IRuntimeFamilyValidator / IFactoryResolver / IRuntimeProfileRepository | Plugin-registry support for later phases. |
Appendix B
Abstract Factory. A pattern for creating families of related objects without binding to concrete classes; the basis of Chapter 4.
Adapter. A concrete class in implementations/ satisfying a contract port.
Composition root. The single place that knows every concrete class and wires them from a profile — DefaultCompositionRoot.
Evidence store. The SQLite database of traces, verdicts, wrapper revisions, proposals, snapshots, and approvals — separate from the MySQL finance DB.
Gap-filler. A stub standing in for a port the real family has not yet implemented (e.g. InMemoryUnitOfWork).
Minion. One of Mazda's five specialist Letta agents, each driving a Claude SDK session.
Port. An ABC in contracts/ describing a capability.
Runtime family. A bundle of ten compatible factories for one profile — IAgentRuntimeFamily.
Stub. A minimal but honest port implementation; some do trivial real work, some raise NotImplementedError.
Verdict. The pass/fail/needs-review judgment of whether the agent actually succeeded.
Wrapper. The versioned, roll-back-able envelope around a fixed LLM (Definition 1.1).
Colophon
Set in the Iowan / Palatino old-style family. This edition supersedes the prior change-log
presentation of the Mazda development status and is maintained alongside the
agent_self_improvement source tree. Identifiers and topology current as of the
Nine-Factory-Families revision.
Operational Addendum
Each nonprofit expense may retain four independent references:
receipt_url for a receipt, document_url for a
bank-downloaded source statement or other primary document,
scanned_statement_url for a paper statement scan, and
moms_ledger for Mom’s composite ledger image.
Supporting-document contract. Attach only the field
corresponding to the current classification and preserve the other three.
Documents may arrive in any order and must converge on one matching
expense rather than create duplicates. An identical reference is an
idempotent success. A different non-empty same-type reference must never
overwrite the original; route it through the existing review-note/log
mechanism and identify the conflict without exposing private paths. Do
not invent a NEEDS_DOCUMENT_VERIFICATION expense status; that
member does not exist in the live enum.
Evidence roots and matching. Reconcile across
readable_documents/receipts/,
readable_documents/bank_statements/, and
readable_documents/scanned_statements/, and
readable_documents/ledger_documents/. A ledger check number
joined to the statement paid date/amount and receipt payee/date/amount is
strong evidence for one expense; amount alone is insufficient. All
attachments go through SupportingDocumentService and its
repository interface, never direct handler SQL.
Legacy classification repair. A repeated incoming scan
associated with multiple unrelated standalone transactions is a source
statement, not a receipt shared by all those expenses. Store it in
document_url. When visual or deterministic classification
proves a legacy reference is in the wrong field, reclassify it through
the supporting-document service/repository boundary so the move is
atomic and unrelated references remain untouched.
The Trainer grades successful tool returns and database-backed evidence
for the selected field, target expense ID, association status, preserved
references, and human-verification requirement. Receipt Only matching
must retain receipt_url while adding
document_url; Mom-ledger reconciliation must attach
moms_ledger to every uniquely matched expense.
Set Category evidence visibility. Every available
evidence type must surface as its own dialog action:
receipt_url produces View Receipt,
document_url produces View Source Document, and
moms_ledger produces View Mom’s Ledger. New and
duplicate-matched statement expenses must persist their statement
reference in document_url; the dashboard may derive a source
from a report directory only to keep legacy records usable. The Trainer
checks every returned expense ID and fails a run when a resolvable stored
reference does not produce its corresponding available action.
Two dashboard/pipeline defects fixed 2026-08-08 — neither was a
Mazda reasoning error, both were environment bugs upstream or downstream
of intake. First: a year-folder archive migration
(migrate_receipts_to_year_folders.py,
migrate_bank_statements_to_year_folders.py in
rol_finances) moved files on disk without ever updating the
matching expenses.source_file/receipt_url rows,
so a correctly-stored reference could silently stop resolving to a real
file months later with no code change on Mazda's side. A manual audit
found and fixed 106 drifted rows across five months (Jan/Feb/Mar 2025 plus
spot-checks of April–Sept 2025, Dec 2024, and Jan 2026); the same
reconciliation is now built into tools/maintenance/
backfill_source_file.py and runs automatically at the end of both
migration scripts, so this should not recur silently again. Second: the
"No high-confidence expense row was found in the image" failure on
check-style/handwritten receipts (a document layout the OCR-only red-box
matcher had no strategy for) is fixed — see the note in the Trainer's own
instructions and mazda_suzuki_escalation_contract.md §4a for
the fix and how it was built (a SwarmForge six-pack run). Both failure
modes were already correctly routed away from Mazda coaching by the
Trainer's existing "Do not coach Mazda... file it as a dashboard defect"
rule — that classification was right the whole time; only the underlying
code has changed.