M
Rol Finances · Autonomous Systems Press
Department of Self-Improving Agents

Mazda

A Self-Improving Letta Agent

First Edition  ·  Revised through the Nine Factory Families  ·  Internal circulation

Mazda · A Developer's ManualPreface

Front Matter

Preface

This manual replaces the change-log that previously occupied this page. A change-log answers the question "what happened last?"; a manual answers the more durable question "how is this thing built, and how do I work on it?" The project has matured past the point where a running diary serves it. What follows is organized the way a course text is organized: foundations first, architecture second, daily practice third, reference material last.

The subject is Mazda — a Letta agent that verifies finance data (bank statements and scanned receipts) inside the larger rol_finances project — together with the framework that surrounds her. That framework has a single, unfashionable thesis, stated here once and defended throughout:

The Central Thesis

We improve the wrapper around a cheap large language model, and we never touch the model itself. The wrapper is everything we can version and roll back: system messages, prompt templates, tool descriptions, context strategy, memory notes, and workflows. The model is a fixed, interchangeable engine. All of our engineering effort — tracing, judging, proposing, A/B testing, gating, activating, rolling back — operates on the wrapper.

A reader who internalizes only that paragraph has the core of the design. The rest of the manual explains the machinery that makes the thesis operational and safe.

How to read this manual

Part I establishes what Mazda is and the live system she runs in. Part II is the heart of the book: the three-layer architecture, the nine factory families that compose a runtime, and the self-improvement loop that closes around them. Part III is operational — how to build, test, and run, and how to verify the live fleet. The appendices are reference catalogs.

Each chapter ends with Exercises. They are not busywork: each one corresponds to a real task a developer on this project will eventually perform. Working them is the fastest way to become productive.

— The maintainers, Autonomous Systems group

· vii ·
Mazda · A Developer's ManualContents

Front Matter

Contents

· ix ·

Part OneFoundations

In which we fix the subject — what Mazda is, what she is not, and the live system in which she runs — before any line of architecture is drawn.

Part I · FoundationsCh. 1 · What Mazda Is

Chapter One

What Mazda Is

Mazda is an orchestrator, not a monolith. She is a single Letta agent who delegates the real work to a small team of specialist agents, each of which drives a Claude Agent SDK session for one slice of the finance-verification problem. Understanding this division of labor is the prerequisite for everything else.

1.1The two faces of the project

It is easy to conflate two things that share the name "Mazda." Keep them distinct:

Definition 1.1 · The Wrapper

The wrapper is the versioned, roll-back-able envelope around a fixed LLM: its system messages, prompt templates, tool descriptions, context-selection strategy, memory notes, and workflow definitions. The framework improves the wrapper; it never fine-tunes or swaps the model as an act of "improvement."

1.2Mazda and her five minions

Mazda is the orchestrator — there is no separate "Orchestrator" agent. She delegates to five minions via Letta's inter-agent messaging. Each minion's specialty is driving a Claude Agent SDK (TypeScript) session for its slice of work. Finance domain logic — totals, vendor key, category, and duplicate verifiers — is exposed to those SDK sessions as MCP tools, not called by Mazda directly.

┌──────────────────────────┐ │ MAZDA │ the orchestrator │ (Letta agent, router) │ delegates · collects · reports └────────────┬─────────────┘ send_letta_message │ ┌──────────────┬───────────┼───────────┬──────────────┐ ▼ ▼ ▼ ▼ ▼ ┌─────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────────────┐ │ router │ │ parser │ │ vendor │ │ receipt │ │categorization│ │ agent │ │ agent │ │ identity │ │ linker │ │ agent │ └────┬────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ └──────┬──────┘ └────────────┴── run_claude_code_sdk ──┴──────────────┘ │ ▼ ┌───────────────────────────────┐ │ Claude-SDK executor :8799 │ spawns a TS Claude session, │ + finance_verifiers_mcp (4) │ which may call MCP verifiers └───────────────────────────────┘
Figure 1.1The delegation topology. One shared run_claude_code_sdk tool fans out to a single executor; the finance verifiers reach the SDK session over MCP.

1.3What Mazda is not

Two earlier designs were considered and deliberately discarded. A new developer will find references to them in old commits and should not revive them:

Caution · Superseded designs

(a) The "give Mazda seven direct finance tools" design, in which the orchestrator called verifiers herself. Superseded by the minion + MCP arrangement above. (b) The separate-TypeScript-package minion path (IClaudeAgentSdkRunner / DefaultMinion / cli.ts) and its Python adapter. The live run_claude_code_sdk tool (one Python tool per agent, inline TS through the :8799 executor) is the canonical Phase 2 implementation; the package path is retired dead code.

1.4Why "self-improving"

The adjective is earned by the control plane, not by any single clever prompt. The framework can observe Mazda doing her job, judge whether she actually succeeded (not merely whether she produced output), propose a concrete edit to her wrapper, run that edit against a baseline behind quality gates, and either activate it or discard it — all with a snapshot it can roll back to. Chapter 5 develops this loop in full. The thesis of the Preface is what keeps it honest: every "improvement" is a wrapper revision, never a model change.

Exercises

  1. (concept) State, in one sentence each, the difference between the runtime fleet and the control-plane framework. Which one imports finance code?
  2. (recall) Name the five minions and the slice of work each owns. (Answer in §2.2.)
  3. (judgment) A teammate proposes adding a fifth finance tool directly to Mazda. Cite the section that tells you why this is the wrong layer, and say where the tool belongs instead.
· 3 ·
Part I · FoundationsCh. 2 · System Topology

Chapter Two

System Topology: The Live Fleet

A framework is only as real as the infrastructure it runs against. This chapter is the field guide to the live deployment: the agents and their identifiers, the executor that backs their SDK sessions, and the finance MCP server they call. Treat the identifiers here as authoritative reference, not as prose to be read once.

2.1The orchestrator

Mazda is the Letta agent agent-6b536cf4-ec88-4290-b595-fed21d14bd8e. She renders six system/ memory blocks into compiled context: persona, human, db_schema, environment, team_agents, and verification_procedure. (How to verify that rendering — the one authoritative method — is §7.2.)

2.2The five minions

All five are git-memory-enabled, each renders system/persona + system/human, and all five share a single run_claude_code_sdk tool — one tool identifier reused across the team, so one client.tools.update covers them all.

Table 2.1The minion roster — live, verified Letta identifiers.
MinionAgent IDSpecialty
mazda-router-agentagent-bc561f63-a5bd-4192-806e-58d92593da2bTask routing
mazda-parser-agentagent-a5063757-46c7-4054-a07d-2b1263db43a8Parsing / structured extraction
mazda-vendor-identity-agentagent-acd624ac-17f2-4a74-aa34-78036cac4d66Vendor normalization
mazda-receipt-linker-agentagent-9a14f800-d848-4914-bfd4-53ab62bc177bReceipt ↔ transaction linking
mazda-categorization-agentagent-c429ff25-c8af-4f1a-a6f1-6d48307e2874Category resolution

Each minion's run_claude_code_sdk(task, context, working_dir) writes an inline TypeScript file that calls claude().withModel('sonnet').allowTools('Read','Write','Edit','Bash','Glob','Grep').inDirectory(workDir).query(prompt).asText() and runs it through the executor. The minions are built by setup_mazda_minions.py on the Windows host.

2.3The Claude-SDK executor

The executor is the small HTTP service that actually spawns a Claude SDK session. It is published at http://100.80.49.10:8799/claude_sdk (fallback 172.17.0.1:8799). It authenticates with a bearer token carried in the tool source; an unauthenticated POST returns 401, an empty body returns 422, and a bare GET returns 405 — the three responses that together prove "alive and validating."

Note · The executor executes, it does not merely validate

A round-trip with the task "reply PONG" returns {"status":"ok","output":"PONG"}. This is the cheap smoke test that confirms a real Claude SDK session ran end-to-end, not just that the request schema was accepted.

2.4The finance MCP server

Finance domain logic reaches the SDK sessions through a FastMCP("finance_verifiers_mcp") server that exposes four deterministic verifier tools over stdio. It is wired into the executor's inline-TS template via .withMCP(...) — covering all five minions at once — and is deployed as a systemd service via deployment/mazda-tools-mcp.service.

Table 2.2The four finance MCP tools.
ToolWhat it checksStatus
verify_statement_totalsLine items sum to the stated totallive
check_vendor_keyVendor key is recognized (vendor_category.yaml)live
check_categoryCategory name is validlive
check_duplicatesRow is not already in the finance DBneeds DB creds

check_duplicates requires pymysql plus finance MySQL credentials; absent those on a given host it returns a structured {"error": …} rather than crashing. The other three are pure Python and always available.

Exercises

  1. (operational) Write the one-line smoke test that proves the executor executes rather than merely validates. What output confirms success?
  2. (reasoning) Three of the four MCP tools are pure Python; one is "best effort." Explain why this asymmetry exists and what it implies for testing on a host without finance DB creds.
  3. (recall) How many client.tools.update calls does it take to update run_claude_code_sdk across all five minions, and why?
· 9 ·

Part TwoThe Architecture

The three layers, the nine factory families that compose a runtime from them, and the self-improvement loop that closes around the whole.

Part II · The ArchitectureCh. 3 · Three Layers

Chapter Three

The Three-Layer Architecture

The framework is built in three layers, strictly ordered by the direction in which dependencies are allowed to point. The rule is simple and it is enforced by tests: program against contracts; concrete code lives behind them. Violating the layering does not merely offend taste — it breaks the test suite's sys.path setup.

3.1The layers

Table 3.1The three layers and what each may depend on.
LayerPackageContentsMay import
Contracts contracts/ Pydantic value objects + ABC ports. The public surface. Only the standard library + Pydantic. Never finance or agent packages.
Implementations implementations/ Concrete adapters, grouped by milestone: persistence, runtime, evaluation, improvement, MCP, reporting, stubs. Contracts; its own siblings. Finance code only by injection.
Agent packages agent_packages/ Per-agent glue — Mazda's task ports, models, and the multi-agent finance orchestrator. Contracts + implementations. This is where agent-specific detail is allowed to live.
Definition 3.1 · Port and Adapter

A port is an abstract base class in contracts/ describing a capability (e.g. ITraceRepository). An adapter is a concrete class in implementations/ that satisfies a port (e.g. SqliteTraceRepository). Callers depend on the port; the choice of adapter is made in exactly one place — the composition root of Chapter 4.

3.2The critical boundary

Caution · The boundary that must never be crossed

The kernel and contracts/ must never import finance packages or any agent package. Mazda is reached only through the generic IAgentPackageFactory and through finance objects injected as Any-typed constructor arguments. The Live*Verifier classes are the model to follow: they wrap real finance services by dependency injection (a VendorCategoryLookup, a DuplicateChecker), never by import.

This is why the evaluation layer ships two verifier families: in-memory stubs (KnownVendorKeyVerifier and friends) for unit tests that touch no infrastructure, and Live*Verifier adapters that wrap the real finance services by injection. The contracts never know which is in use.

3.3How new behavior enters the system

The convention is invariant across the whole project: new behavior starts as a contract, then gets a concrete implementation behind it. A Pydantic model or an ABC goes into contracts/ first; the adapter follows in implementations/; commonly used names are re-exported from the relevant __init__.py. This discipline is what makes the factory families of the next chapter possible — you cannot manufacture what you have not first specified.

Exercises

  1. (rule) A pull request adds from nonprofit_finance import ... to a file in contracts/. State which rule it breaks and what mechanism will catch it.
  2. (design) You must add a new capability "redact PII from a trace." In which layer does the ABC go, in which the adapter, and which file re-exports the name?
  3. (recall) Why does the evaluation layer ship two verifier families instead of one?
· 17 ·
Part II · The ArchitectureCh. 4 · Nine Factory Families

Chapter Four

The Nine Factory Families

This is the chapter the architecture is built around. The framework composes a runtime from families of related objects using the Abstract Factory pattern. Nine factories cover the nine subsystems; a tenth — the agent-package factory — joins them inside a single runtime family. A composition root wires the family from a profile, and an agent kernel runs tasks through it. By the end of this chapter you will be able to add a backend, swap a whole family, or trace one task from request to persisted verdict.

4.1Why Abstract Factory

Motivation 4.1

The kernel must create many related objects — a trace repository, a verdict judge, an event bus — without binding to their concrete classes. The Abstract Factory pattern lets us swap a whole compatible family (say, SQLite persistence + a cheap-LLM client + finance evaluation) without the kernel ever changing. The cost is one extra layer of indirection; the benefit is that "which backend" is a decision made in exactly one place.

Each factory is an ABC in contracts/factories.py; each create_* method returns a port, never a concrete type. The full method-by-method catalog is Appendix A; the families themselves are Table 4.1.

Table 4.1The nine factory families (plus the agent-package factory), their contracts, and concrete implementations.
#Factory (ABC)ManufacturesConcrete implementation
1IPersistenceFactoryschema manager, unit of work, 5 repositoriesSqlitePersistenceFactory(db_path) real
2ILlmFactoryclient, request builder, parser, token counter, budget guardStubLlmFactory(model_name) stub
3IPromptFactorysystem-message provider/builder, template provider, composer, validator, diffStubPromptFactory stub
4IToolFactoryregistry, adapter, permission policy, execution loggerStubToolFactory stub
5IContextFactoryselector, budgeter, pack builder, memory/note readers + writers, staleness detectorStubContextFactory stub
6IWorkflowFactoryworkflow runnerStubWorkflowFactory stub
7IEvaluationFactoryevaluator, verdict judge, classifier, regression detector, gate chain, 4 finance verifiersFinanceEvaluationFactory(...) real
8IImprovementFactoryproposal generator, experiment runner, approval gateway, snapshot store, activation + rollbackDefaultImprovementFactory(...) real
9IReportingFactoryevent bus, trace report builder, dashboard builder, regression report builderDefaultReportingFactory real
+IAgentPackageFactoryMazda / Scissari / Frita packages by nameDefaultAgentPackageFactory real

4.2Real families and stub families

Not every subsystem has a production backend yet, and the architecture does not pretend otherwise. Four families — persistence, evaluation, improvement, reporting — have real implementations wired to SQLite and the finance verifiers. Five families — LLM, prompt, tool, context, workflow — currently ship stubs in implementations/stubs/.

Definition 4.2 · Stub vs. Gap-filler

A stub is a minimal but honest port implementation. Some stubs do real, trivial work (the workflow runner genuinely iterates steps in sequence; the request builder genuinely assembles a request); others raise NotImplementedError for operations that need infrastructure we have not built (an LLM complete(), a memory write()). A gap-filler is a stub that stands in for a port the real family hasn't implemented yet — e.g. InMemoryUnitOfWork and InMemoryEvaluationRepository, used by the SQLite persistence factory for its two not-yet-persisted ports.

This honesty matters at runtime: the kernel expects some ports to raise NotImplementedError and degrades gracefully (§4.6). Stubs are not technical debt to be hidden; they are placeholders with a precise contract, and they make the whole family composable today.

4.3The runtime family — bundling ten factories

A single profile uses one factory from each family. The IAgentRuntimeFamily port bundles all ten as read-only properties, so a subsystem can be swapped by swapping one property's factory. The concrete bundle is DefaultAgentRuntimeFamily; it takes the ten factories as keyword arguments and exposes them.

contracts/runtime.py — the bundle (abridged)
class IAgentRuntimeFamily(ABC):
    """A bundle of compatible factories for one runtime profile."""
    @property
    @abstractmethod
    def persistence_factory(self)  -> IPersistenceFactory: ...
    @property
    @abstractmethod
    def llm_factory(self)          -> ILlmFactory: ...
    # ... prompt, tool, context, workflow, evaluation,
    #     improvement, reporting, agent_package ...

4.4The composition root — the only place that knows concrete classes

If the kernel must never name a concrete class, something must. That something is the composition root. It is the single seam where abstraction is traded for concreteness, and it is deliberately the only such seam in the system.

Definition 4.3 · Composition Root

DefaultCompositionRoot.build_kernel(profile) reads a RuntimeProfile and instantiates every concrete factory, bundles them into a DefaultAgentRuntimeFamily, and returns a DefaultAgentKernel. It is the only file in the framework permitted to import concrete factory classes.

The wiring order is not arbitrary — it encodes the cross-family dependencies. Persistence is built first (and its schema created) because the improvement factory needs persistence's wrapper-revision and improvement repositories injected into it. Stubs come next, evaluation is configured from profile.extra, then improvement is cross-wired, and reporting last.

implementations/runtime/composition_root.py (abridged)
class DefaultCompositionRoot(ICompositionRoot):
    def build_kernel(self, profile: RuntimeProfile) -> IAgentKernel:
        db_path = profile.db_path or "agent_improvement.sqlite3"

        persistence = SqlitePersistenceFactory(db_path=db_path)
        persistence.create_schema_manager().create_schema()   # schema first

        llm      = StubLlmFactory(model_name=profile.llm)
        prompt   = StubPromptFactory()
        tool     = StubToolFactory()
        context  = StubContextFactory()
        workflow = StubWorkflowFactory()

        evaluation = FinanceEvaluationFactory(            # configured from the profile
            known_vendor_keys=set(profile.extra.get("known_vendor_keys", [])),
            vendor_category_map=profile.extra.get("vendor_category_map", {}),
            ...)

        agent_package = DefaultAgentPackageFactory()

        improvement = DefaultImprovementFactory(          # cross-wired with persistence
            db_path=db_path,
            wrapper_repo=persistence.create_wrapper_revision_repository(),
            improvement_repo=persistence.create_improvement_repository())

        reporting = DefaultReportingFactory()

        family = DefaultAgentRuntimeFamily(
            persistence=persistence, llm=llm, prompt=prompt, tool=tool,
            context=context, workflow=workflow, evaluation=evaluation,
            improvement=improvement, reporting=reporting,
            agent_package=agent_package)

        return DefaultAgentKernel(family=family, profile=profile)

The RuntimeProfile is a pure data record — it names the families ("sqlite", "gemini_flash_lite", "mazda", "finance_rules") and carries an extra dictionary for evaluation configuration and a db_path. The canonical first profile is sqlite_gemini_flash_lite_mazda_default. Selecting a different backend is, by design, an edit to a profile — not to the kernel.

4.5The agent kernel — one task, end to end

The kernel depends only on IAgentRuntimeFamily. It never learns whether persistence is SQLite, whether the LLM is one model or another, or who Mazda is. Its run_task is short enough to read in full, and reading it is the best way to see the families cooperate.

RunRequest │ ▼ agent_package_factory.create_agent_package(name).runner() ┌──────────────┐ │ runner.run │ ─────────────► RunResult (output, tool_calls, tokens) └──────┬───────┘ ▼ assemble TraceRecord ──────► persistence_factory.create_trace_repository().save_trace() │ ▼ evaluation_factory.create_verdict_judge().judge(trace) VerdictRecord ────► save_verdict() [degrades on NotImplementedError] │ ▼ reporting_factory.create_event_bus().publish(...) RunOutcome(trace, verdict)
Figure 4.1One pass of DefaultAgentKernel.run_task. Four of the ten families participate; the verdict and event steps degrade gracefully when a stub raises NotImplementedError.
Note · Graceful degradation is a feature

The kernel wraps the verdict and event-publishing steps in try / except NotImplementedError. This is what lets a runtime composed of real persistence but stubbed reporting still produce a valid RunOutcome. The trace is always saved; the verdict is saved when an evaluation backend exists. The composition stays whole even while half its families are stubs.

4.6The families are additive, not a rewrite

A crucial property for anyone maintaining existing code: the factory layer was added on top of the direct-construction code, not in place of it. Every entry point and every pre-existing test that wires objects by hand continues to work unchanged. The factories are a second, optional path to the same concrete classes — proven by a dedicated test that constructs objects both ways and asserts they are equivalent.

4.7Worked example: adding a MySQL persistence family

Suppose the evidence store must move from SQLite to MySQL. The Abstract Factory pattern makes the blast radius precise:

  1. Write MySqlTraceRepository and siblings as adapters behind the existing persistence ports.
  2. Write MySqlPersistenceFactory(IPersistenceFactory) returning them.
  3. In the composition root, branch on profile.persistence == "mysql" to choose the factory.

The kernel, the runtime family, every evaluator, and every test that depends on the persistence port are untouched. That containment is the entire return on the indirection the pattern costs.

Exercises

  1. (trace) Follow a single RunRequest through run_task and list, in order, every factory whose create_* method is called. (Answer: see Figure 4.1.)
  2. (design) Why must persistence be constructed before improvement in the composition root? Name the two objects that flow from one to the other.
  3. (implementation) Sketch the three steps to add a GeminiLlmFactory that replaces the stub. Which existing files change, and which do not?
  4. (judgment) The workflow factory's runner is a stub, yet it is "real" enough to use. Reconcile this with Definition 4.2.
  5. (verification) Explain how the test suite proves the factories are additive rather than a replacement.
· 25 ·
Part II · The ArchitectureCh. 5 · The Loop

Chapter Five

The Self-Improvement Loop

With a runtime composed and a kernel able to run one task, we can close the loop that gives the project its name. The loop runs the agent, judges whether it truly succeeded, proposes a wrapper edit, A/B tests that edit behind gates, and activates the winner with a snapshot it can roll back to. It imports no finance code and names no agent; it reaches the agent only through ports.

5.1The six stations

┌──────────┐ ┌──────────┐ ┌────────────┐ ┌────────────┐ ┌────────┐ ┌────────────┐ │ RUN │──►│ JUDGE │──►│ PROPOSE │──►│ EXPERIMENT │──►│ GATE │──►│ ACTIVATE │ │ (trace) │ │ (verdict)│ │ (edit) │ │ (A vs. B) │ │ (4×) │ │ / ROLLBACK │ └──────────┘ └──────────┘ └────────────┘ └────────────┘ └────────┘ └────────────┘ │ │ │ └──────────────┴─────────────── evidence store (SQLite) ───────────────────┘
Figure 5.1The six stations of the loop, all writing to one evidence store.

1 · Run & trace

The kernel runs a task and persists a TraceRecord: inputs, outputs, tool calls, token usage, and the four wrapper revisions in force. This is the raw evidence everything downstream reasons over.

2 · Judge

Judging answers "did the agent actually succeed?" — not merely "did it return text?" It is deterministic-rules-first: the finance verifiers produce structured evidence, a rule-based evaluator scores it into a ScoreCard, and the verdict judge assigns pass / fail / needs-review. An LLM judge strategy is consulted only for genuinely ambiguous cases (an unrecognized vendor, a fuzzy duplicate). A non-passing verdict is then labeled with a finance FailureType.

Principle 5.1 · Determinism before judgment

Cheap, deterministic rules decide every case they can. The expensive, non-deterministic LLM judge is a fallback for the residue of ambiguity — never the first resort. This keeps verdicts reproducible and keeps cost down.

3 · Propose

From failure patterns, a proposal generator suggests a concrete wrapper edit. Edits are modeled as commands: PromptPatchCommand, ToolDescriptionPatchCommand, MemoryNoteCommand, ContextRulePatchCommand. A command is a reified, inspectable, reversible change — exactly the granularity the gates and rollback need.

4 · Experiment

A proposed edit is not trusted on its say-so. A baseline-versus-candidate experiment runs both wrappers and a regression detector compares their scorecards, so an "improvement" that quietly regresses some other case is caught before it ships.

5 · Gate

The candidate must clear a GateChain of four independent gates before it is eligible for activation:

Table 5.1The four promotion gates.
GateQuestion it asks
SafetyGateDoes the change violate a safety invariant?
CostGateDoes it blow the token / latency budget?
RegressionGateDid any previously-passing case regress?
UsefulnessGateIs the improvement actually material?

6 · Activate & roll back

A gated winner is activated through a human approval gateway, with a snapshot taken first. If the change misbehaves in practice, the rollback service restores the known-good wrapper revision by id. Approval and activation records land in the same evidence store as everything else, so the whole history is auditable.

5.2Where the loop meets Mazda

The loop is finance-blind infrastructure; only the thing being wrapped is Mazda-specific. The mapping is direct:

Table 5.2How the generic loop applies to the Mazda fleet.
Loop conceptMazda usage
IAgentRunner / run_taskRun Mazda's orchestration step; collect each minion's SDK-session result.
Trace / evidence storeRecord Mazda→minion delegations and each SDK session (inputs, outputs, tool calls, tokens).
IEvaluator familyOrchestration verdict; the finance verifiers become one pluggable evaluator.
WrapperRevisionVersioned delegation logic + per-minion SDK task templates (prompts, allowed tools, params).
Proposal / Experiment / Gate / RollbackPropose, A/B, gate, and human-approve a change to delegation strategy or a minion template — rollback-safe.

Exercises

  1. (principle) State Principle 5.1 in your own words and give one reason cost, and one reason reproducibility, both favor it.
  2. (design) Why are wrapper edits modeled as command objects rather than as direct mutations? Connect your answer to the gate and rollback stations.
  3. (reasoning) A candidate improves the failing case but trips the RegressionGate. What happened, and what does the loop do next?
  4. (mapping) Give the Mazda-specific meaning of WrapperRevision in this project.
· 43 ·

Part ThreePractice

From architecture to the keyboard — how to build, test, and run the framework, and how to verify the live fleet without re-diagnosing solved problems.

Part III · PracticeCh. 6 · Build & Test

Chapter Six

Building, Testing, and Running

This chapter is the lab manual. It assumes you are working in tools/self_improving_agent/ and that the virtual environment lives at the repository root in ../../.venv, not inside this subdirectory — a detail that trips newcomers and is worth memorizing.

6.1Running the tests

Use .venv/bin/python explicitly; the system python may be absent and python3 may lack pytest. The conftest.py + pytest.ini pair makes the package importable regardless of which directory you invoke from.

# everything (live integration tests self-skip if env unset)
../../.venv/bin/python -m pytest

# one file — e.g. the factory-family suite
../../.venv/bin/python -m pytest agent_self_improvement/tests/test_phase_02b_factory_families.py

# one test by name substring
../../.venv/bin/python -m pytest -k "vendor_key"

# unit only / certification-blocking integration only
../../.venv/bin/python -m pytest -m "not integration"
../../.venv/bin/python -m pytest -m p0
Note · Two kinds of test

Unit tests (tests/test_phase_NN_*.py) use duck-typed fakes — no MySQL, no Letta, no network — and should always pass. Integration tests (tests/integration/p0|p1|p2) are marked and skip themselves unless the right environment variables are set. The severity tiers are p0 (blocks certification), p1 (blocks multi-agent expansion), and p2 (hardening).

6.2The test phases as a learning map

The phase-numbered test files double as a curriculum: each maps to a milestone of the architecture. The factory work of Chapter 4 lives in its own phase file.

Table 6.1Selected test phases and the chapter each illustrates.
Test fileCoversChapter
test_phase_00_value_objects.pyPydantic value objects, typed ids3
test_phase_01_database_contracts.pySQLite evidence store3, 5
test_phase_02_wrapper_bundle.pyWrapper revisions + snapshot/rollback5
test_phase_02b_factory_families.pyThe nine factories, runtime family, composition root, kernel (18 tests)4
test_phase_03_agent_runner.pyAgent runner port + fakes4
test_phase_04_verdicts.pyVerifiers → scorecard → verdict → failure type5
test_phase_05–07Proposals, experiments, activation + reports5
test_phase_08–09Approval audit, concurrency guard5

Across the suite there are over two hundred test functions; the factory-family file alone contributes eighteen, proving that every factory creates its ports, the runtime family bundles them, the composition root builds a kernel, the kernel runs a task with a persisted trace, and direct instantiation still works.

6.3Entry points

Receipt verification demo — mazda_verify_receipt.py

Runs Mazda's deterministic verifiers over one parsed receipt JSON, optionally asks the live Letta agent for a narrative, persists a trace + verdict to SQLite, and writes an HTML report to receipt_runs/.

python3 mazda_verify_receipt.py -f walgreens_2025-01-22.json          # full run
python3 mazda_verify_receipt.py -f receipt.json --no-letta            # skip live Letta turn
python3 mazda_verify_receipt.py -f receipt.json --db /tmp/run.sqlite3 --output out.html

Team orchestrator — mazda_run_team.py

../../.venv/bin/python mazda_run_team.py --check-ready                  # verify specialist IDs
../../.venv/bin/python mazda_run_team.py --dry-run  --source-uri /tmp/test.pdf  # offline workflow
../../.venv/bin/python mazda_run_team.py --live     --source-uri /tmp/test.pdf  # live specialists

Finance MCP server

implementations/finance_mcp/server.py exposes the four verifier tools over stdio; it is deployed as a systemd service via deployment/mazda-tools-mcp.service behind mcp-proxy on port 8791.

Exercises

  1. (setup) From tools/self_improving_agent/, write the exact command to run only the factory-family suite. Why must you spell out the venv path?
  2. (operational) You set no integration environment variables and run the full suite. What happens to the p0 tests, and why is that correct behavior rather than a failure?
  3. (exploration) Run the receipt demo with --no-letta. Which stations of the Chapter 5 loop still execute, and which are skipped?
· 55 ·
Part III · PracticeCh. 7 · Operating the Fleet

Chapter Seven

Operating the Live Fleet

The live fleet has a small number of recurring operational questions and exactly one correct way to answer the most important of them. This chapter records those answers so that a new instance never re-diagnoses a problem the project has already solved.

7.1Reaching the live Letta server

7.2The one correct way to verify memory

The single authoritative test of whether an agent renders its memory is the <projection> count from a recompile. Everything else is a red herring.

curl -s -X POST http://100.80.49.10:8283/v1/agents/<AGENT_ID>/recompile \
  | grep -oE '<projection>[^<]*</projection>'
# Each line = one system/ block rendered into context.
# Mazda = 6 · each minion = 2 · zero = amnesiac.
Caution · Do not re-diagnose memory from the wrong signal

The agents.system DB column is the base prompt template — block content is injected at compile time and is never stored there, so it will always look "empty" of memory. An empty memfs repo (zero commits) is likewise fine: on this deployment the system/-labeled block drives rendering. And /v1/agents/<id>/context's system_prompt field is stale for dormant agents. Trust only the recompile projection count.

Note · A timing quirk

A recompile reads freshly-attached blocks on the next call, not the same instant they attach. If the first projection grep returns empty right after attaching a block, re-run it once.

7.3Two operational traps worth memorizing

Single-file bind mounts

Editing executor_server.py in place changes its inode, which breaks docker-desktop's single-file bind mount and makes docker restart fail with "no such file or directory." Recovery is to recreate via the idempotent deploy_frita_executor.sh. Directory mounts do not have this problem.

Root + IS_SANDBOX

The executor runs as root, so claude --dangerously-skip-permissions is refused unless IS_SANDBOX=1 is set. It is now exported both in the SDK subprocess environment and as a container variable, so .skipPermissions() works under root.

7.4Required reading, in order

  1. This manual — Parts I–III for current state and architecture.
  2. memory_system_plan.html — the memory policy: author memory as system/** memfs files; blocks are a read-only projection; never raw POST/PATCH /v1/blocks.
  3. agent_self_improvement_plan.html — the design source for the control-plane loop and the nine factory families. (Its older "Mazda gets seven finance tools" framing is superseded; see §1.3.)

Exercises

  1. (operational) An agent appears "amnesiac." List the three wrong signals you must not use to confirm it, and the one right one you must.
  2. (recovery) docker restart frita-executor fails with "no such file or directory." What did someone just do, and what is the recovery command?
  3. (reasoning) Why does the executor need IS_SANDBOX=1? What would silently fail without it?
· 63 ·
Back MatterAppendix A

Appendix A

Interface & Factory Catalog

A method-by-method reference for the ten factories of Chapter 4. Every method returns a port (an ABC), never a concrete class.

Table A.1contracts/factories.py — the create methods.
Factorycreate_* methods
IPersistenceFactoryschema_manager · unit_of_work · trace_repository · wrapper_revision_repository · prompt_artifact_repository · tool_artifact_repository · evaluation_repository · improvement_repository
ILlmFactoryclient · request_builder · response_parser · token_counter · budget_guard
IPromptFactorysystem_message_provider · system_message_builder · template_provider · composer · validator · diff
IToolFactoryregistry · adapter · permission_policy · execution_logger
IContextFactoryselector · budgeter · pack_builder · memory_reader · memory_writer · note_reader · note_writer · staleness_detector
IWorkflowFactoryrunner
IEvaluationFactoryevaluator · verdict_judge · failure_classifier · regression_detector · gate_chain · totals_verifier · vendor_key_verifier · category_verifier · duplicate_verifier
IImprovementFactoryproposal_generator · experiment_runner · approval_gateway · snapshot_store · activation_service · rollback_service
IReportingFactoryevent_bus · trace_report_builder · dashboard_builder · regression_report_builder
IAgentPackageFactorycreate_agent_package(name) · supported_agents()
Table A.2The runtime-assembly contracts in contracts/runtime.py.
ContractRole
RuntimeProfileData-only record naming each family + db_path + extra.
IAgentRuntimeFamilyBundles all ten factories as read-only properties.
ICompositionRootbuild_kernel(profile) → IAgentKernel; only place that names concrete classes.
IAgentKernelrun_task(request) → RunOutcome.
RunOutcome{ trace, verdict? } — the product of one run.
IRuntimeFamilyValidator / IFactoryResolver / IRuntimeProfileRepositoryPlugin-registry support for later phases.
· 71 ·
Back MatterAppendix B

Appendix B

Glossary

Abstract Factory. A pattern for creating families of related objects without binding to concrete classes; the basis of Chapter 4.

Adapter. A concrete class in implementations/ satisfying a contract port.

Composition root. The single place that knows every concrete class and wires them from a profile — DefaultCompositionRoot.

Evidence store. The SQLite database of traces, verdicts, wrapper revisions, proposals, snapshots, and approvals — separate from the MySQL finance DB.

Gap-filler. A stub standing in for a port the real family has not yet implemented (e.g. InMemoryUnitOfWork).

Minion. One of Mazda's five specialist Letta agents, each driving a Claude SDK session.

Port. An ABC in contracts/ describing a capability.

Runtime family. A bundle of ten compatible factories for one profile — IAgentRuntimeFamily.

Stub. A minimal but honest port implementation; some do trivial real work, some raise NotImplementedError.

Verdict. The pass/fail/needs-review judgment of whether the agent actually succeeded.

Wrapper. The versioned, roll-back-able envelope around a fixed LLM (Definition 1.1).

Colophon

Set in the Iowan / Palatino old-style family. This edition supersedes the prior change-log presentation of the Mazda development status and is maintained alongside the agent_self_improvement source tree. Identifiers and topology current as of the Nine-Factory-Families revision.

· 76 ·
Document IntakeSupporting Evidence

Operational Addendum

Four Independent Supporting Documents

Each nonprofit expense may retain four independent references: receipt_url for a receipt, document_url for a bank-downloaded source statement or other primary document, scanned_statement_url for a paper statement scan, and moms_ledger for Mom’s composite ledger image.

Supporting-document contract. Attach only the field corresponding to the current classification and preserve the other three. Documents may arrive in any order and must converge on one matching expense rather than create duplicates. An identical reference is an idempotent success. A different non-empty same-type reference must never overwrite the original; route it through the existing review-note/log mechanism and identify the conflict without exposing private paths. Do not invent a NEEDS_DOCUMENT_VERIFICATION expense status; that member does not exist in the live enum.

Evidence roots and matching. Reconcile across readable_documents/receipts/, readable_documents/bank_statements/, and readable_documents/scanned_statements/, and readable_documents/ledger_documents/. A ledger check number joined to the statement paid date/amount and receipt payee/date/amount is strong evidence for one expense; amount alone is insufficient. All attachments go through SupportingDocumentService and its repository interface, never direct handler SQL.

Legacy classification repair. A repeated incoming scan associated with multiple unrelated standalone transactions is a source statement, not a receipt shared by all those expenses. Store it in document_url. When visual or deterministic classification proves a legacy reference is in the wrong field, reclassify it through the supporting-document service/repository boundary so the move is atomic and unrelated references remain untouched.

The Trainer grades successful tool returns and database-backed evidence for the selected field, target expense ID, association status, preserved references, and human-verification requirement. Receipt Only matching must retain receipt_url while adding document_url; Mom-ledger reconciliation must attach moms_ledger to every uniquely matched expense.

Set Category evidence visibility. Every available evidence type must surface as its own dialog action: receipt_url produces View Receipt, document_url produces View Source Document, and moms_ledger produces View Mom’s Ledger. New and duplicate-matched statement expenses must persist their statement reference in document_url; the dashboard may derive a source from a report directory only to keep legacy records usable. The Trainer checks every returned expense ID and fails a run when a resolvable stored reference does not produce its corresponding available action.

Two dashboard/pipeline defects fixed 2026-08-08 — neither was a Mazda reasoning error, both were environment bugs upstream or downstream of intake. First: a year-folder archive migration (migrate_receipts_to_year_folders.py, migrate_bank_statements_to_year_folders.py in rol_finances) moved files on disk without ever updating the matching expenses.source_file/receipt_url rows, so a correctly-stored reference could silently stop resolving to a real file months later with no code change on Mazda's side. A manual audit found and fixed 106 drifted rows across five months (Jan/Feb/Mar 2025 plus spot-checks of April–Sept 2025, Dec 2024, and Jan 2026); the same reconciliation is now built into tools/maintenance/ backfill_source_file.py and runs automatically at the end of both migration scripts, so this should not recur silently again. Second: the "No high-confidence expense row was found in the image" failure on check-style/handwritten receipts (a document layout the OCR-only red-box matcher had no strategy for) is fixed — see the note in the Trainer's own instructions and mazda_suzuki_escalation_contract.md §4a for the fix and how it was built (a SwarmForge six-pack run). Both failure modes were already correctly routed away from Mazda coaching by the Trainer's existing "Do not coach Mazda... file it as a dashboard defect" rule — that classification was right the whole time; only the underlying code has changed.

· 77 ·