Self-Improvement Systems · Comparative Review
A review of Mazda's production self-improvement control plane against the Stanford IRIS Lab Meta-Harness framework (arXiv:2603.28052, github.com/stanford-iris-lab/meta-harness), with concrete, gate-preserving adoption recommendations.
Mazda and Meta-Harness solve two different halves of the same problem, and each is strong exactly where the other is weak.
Mazda is a production control plane: a ports-and-adapters framework (nine factory families,
six-station loop) that runs an agent, judges it deterministically, proposes a wrapper edit, A/B tests it,
gates it (safety / cost / regression / usefulness), and activates it behind human approval with snapshots
and rollback. Its governance is far ahead of Meta-Harness. But its generative half is thin: the
live MazdaProposalGenerator is a hand-written lookup table mapping six FailureType
values to six canned English patch strings. Mazda can only ever “propose” what a developer
pre-wrote into that table. It also has no validation/test split, no cross-iteration search history, and no
autonomous multi-iteration loop.
Meta-Harness is a research search engine: an outer loop where a full coding agent (Claude Code via
claude_wrapper.py) reads scores and raw execution traces, forms falsifiable hypotheses,
prototypes mechanisms, and writes arbitrary new harness code; an inner loop scores each candidate on a
search set; a Pareto frontier and append-only evolution_summary.jsonl accumulate evidence across
iterations; and a hard firewall separates validation (used during evolution) from held-out test (touched
only once, at an explicit, irreversible finalization step). It has essentially no production governance:
no gates beyond score comparison, no human approval, no activation/rollback, and it warns that candidates
cost real money (~$500/iteration for TB2).
ProposalCommand objects into Mazda's existing gate
chain, plus a validation/test split, a frontier + evolution summary, and anti-leakage source checks.
Nothing downstream of Propose needs to change, which is precisely what Mazda's interface-first design was
built for.
| Dimension | Mazda (agent_self_improvement) | Meta-Harness |
|---|---|---|
| What it improves | The wrapper around a fixed LLM: prompts, tool descriptions, context strategy, memory notes, workflows — for one live finance fleet | The harness: arbitrary Python code around a fixed model (memory systems, agent scaffolds) — per domain, offline |
| Candidate generation | Deterministic: FailureType → pre-written patch text (proposals.py); FailurePatternSuggester groups repeats (≥2) into one proposal |
LLM proposer agent (Claude Code, max effort) reads history + traces, forms 1–3 falsifiable hypotheses per iteration, mandatorily prototypes, writes new code |
| Search space | Fixed menu of 4 command types over existing artifacts | Arbitrary code; skill prompts forbid parameter-only variants and demand new mechanisms |
| Use of traces | Traces persisted to SQLite; consumed by verifiers and pattern counting; not deeply read by any proposer | Proposer deep-reads raw trajectories / log.jsonl as the primary analysis step; proposer's own sessions logged to experience/ (tools, files read, tokens, cost) |
| Historical evidence | Evidence store holds traces / verdicts / proposals / approvals — audit-oriented, per-event | evolution_summary.jsonl (per-candidate hypothesis, axis, delta, outcome, timing), frontier_val.json (best per dataset/task), per-iteration post-mortem reports — search-oriented, cross-iteration |
| Evaluation | Deterministic-rules-first verifiers → ScoreCard → verdict; LLM judge only for ambiguity | Task metric (accuracy / pass rate) on a search set; noise handled by multi-trial evals |
| Held-out test isolation | Absent. A/B runs on task_inputs with no split; nothing prevents optimizing against the same cases forever |
Core design. Evolution sees only val; test results live in a directory never exposed; --test finalization freezes the run permanently (finalized.json blocks further evolution) |
| Overfitting prevention | Only implicit (RegressionGate on baseline cases) | Val/test split, anti-overfitting skill rules (no task names in code), universal + per-task forbidden_references denylist checked against candidate source before spending eval budget |
| Gates | 4-gate chain: Safety, Cost, Regression, Usefulness — explicit contracts | None as such; frontier selection is the only “gate”; TB2 loop does track cost/token rollout_metrics per candidate |
| Human approval / activation / rollback | Approval gateway, snapshot store, activation + rollback services, audit records | None. Winner = frontier top; adoption is manual / out-of-band |
| Trust boundary | Layering rules (contracts import nothing concrete); generated code never exists — proposals are text patches | Generated code is arbitrary Python, never imported by the controller; run in a short-lived sandboxed subprocess (Harbor pilot) |
| Autonomy | One proposal per failing trace; no autonomous multi-iteration loop shipped | Fully autonomous N-iteration loop with resume, fresh-start, interrupt handling |
| Cost accounting | CostGate, per-run token usage in traces, estimated_cost_usd in experiments |
Per-candidate propose/bench/wall timing, tokens, $cost_usd per proposer session; explicit budget discipline (bring-up ladder: 1 task → 30 → 89) |
| Reproducibility | SQLite evidence store, wrapper revisions by id, 200+ provider-free tests | Run-name-isolated log dirs, full raw event logs, import-validation before eval, provider-free test suites |
| Maturity | 4 of 9 factory families real; LLM/prompt/tool/context/workflow are stubs; loop not yet closed end-to-end in production | “Cleaned-up paper code, not tested beyond verifying it runs”; research-grade |
meta_harness.py: FINALIZED state machine, test results in a directory the proposer never sees).evolution_summary.jsonl — one JSON row per candidate with hypothesis, axis, delta vs. frontier, outcome, timing — plus frontier_val.json.claude_wrapper.py's experience/ layout: meta.json, response.md, events.jsonl, per-tool-call files). Drop-in; it's a standalone file.validate_candidates and TB2's validate_agent_class).LlmProposalGenerator adapter behind the existing IProposalGenerator port, driven by a claude_wrapper-style subprocess with a Mazda-specific SKILL.md prior, emitting the same ImprovementProposal/command objects. Everything downstream (experiment, gates, approval, rollback) is untouched.Rationale.expected_benefit, which exists but is canned text today.forbidden_references source check (Harbor controller.py: validate_source + universal denylist). Adapt as a fifth gate or a SafetyGate extension: proposed patch text must not reference held-out case identifiers, specific vendor names from eval fixtures, verifier internals, etc.rollout_metrics. Mazda's CostGate exists; the ladder is an evaluation-ordering policy to bolt onto the experiment runner.BaselineCandidateExperimentRunner + ScorecardRegressionDetector are equivalent to (and cleaner than) Meta-Harness's frontier delta.estimated_cost_usd, CostGate.--fresh wipe-and-restart semantics — conflicts with Mazda's append-only audit evidence.Each item: idea → gap → integration point → benefit / risks → required evidence → priority & effort.
Gap: Mazda has no held-out isolation; repeated gate-passing against the same cases will silently overfit the wrapper to its own regression suite.
Integration: Split the evaluation case corpus (receipt fixtures, statement cases) into search and holdout sets at the persistence layer; ABRunConfig.task_inputs draws only from search. Add a FinalizationService in the improvement family: holdout is evaluated only when a candidate has passed all four gates and is queued for human approval, and a finalized flag on the wrapper-revision record blocks further tuning against that holdout snapshot.
Benefit: The UsefulnessGate's answer becomes trustworthy; approval decisions rest on unseen data.
Risks: Mazda's case corpus is small — a split shrinks statistical power. Mitigate with periodic holdout rotation (rotate, then retire the old holdout into search).
results/ dir); a demonstration that at least one past “improvement” scores differently on the split.IProposalGeneratorGap: the deterministic proposer can only re-emit six pre-written sentences; it cannot learn from traces or invent a fix for an unanticipated failure (everything else falls into the MEMORY_NOTE catch-all).
Integration: new adapter implementations/improvement/llm_proposals.py::ClaudeProposalGenerator(IProposalGenerator) that (a) pulls the failing trace cluster + recent verdicts + current wrapper artifacts from the evidence store, (b) invokes a vendored claude_wrapper.run() with a Mazda SKILL.md prior, (c) parses a structured pending_proposal.json into the existing ImprovementProposal/ChangeSet command objects, (d) falls back to MazdaProposalGenerator on parse failure. Selected in the composition root via profile.improvement == "llm_proposer" — a profile edit, not a kernel change.
Benefit: proposals become trace-grounded and open-ended while remaining inspectable text patches that flow through the unchanged experiment/gate/approval/rollback pipeline. This is Meta-Harness's engine inside Mazda's brakes.
Risks: non-determinism in proposal content (acceptable: verdicts stay deterministic; the proposal is gated); prompt-injection via trace content into the proposer (constrain proposer tools to Read-only over an exported evidence snapshot, never live systems); token cost (log per-session cost à la claude_wrapper, enforce via CostGate).
Gap: the evidence store records events but nothing summarizes what has been tried and with what delta, so a proposer (human or LLM) can't avoid re-proposing failed ideas.
Integration: an evolution_summary table (or JSONL exported by the reporting family) with per-proposal hypothesis, target artifact, gate outcomes, val delta; a frontier view of best-known wrapper revision per task type. Feed both into the P1 proposer's context.
Benefit: cross-iteration learning; the documented Meta-Harness failure mode (“parameter variants regress or tie”) is avoided by making history visible.
Gap: nothing stops a proposal (especially an LLM one) from embedding holdout-case specifics (“Walgreens receipts always…”) — overfitting by memorization.
Integration: LeakageGate in implementations/evaluation/gates.py, configured with a denylist derived from holdout fixtures (vendor names, filenames, case ids) + a universal list, string-checked against patch_text — a direct port of controller.py::validate_source. Add it to the GateChain (making it five gates).
Benefit: cheap, deterministic, runs before any eval spend.
Risks: false positives on legitimately common substrings; keep the list reviewable in config.
experience/)Gap: if P1 ships, the proposer's own reasoning process is unaudited — contrary to Mazda's auditability principle.
Integration: vendor claude_wrapper.py as-is; point log_dir at an experience/ sibling of the SQLite store; store the session directory path on the ImprovementProposal record.
Benefit: every activated change traces back to the exact proposer session (tools used, files read, tokens, cost) — extends Mazda's audit chain upstream into generation.
Gap: experiment outcomes aren't distilled into takeaways; expected_benefit is boilerplate, never scored against reality.
Integration: add hypothesis and prediction to Rationale; after each experiment, the reporting family (or the P1 proposer's next session, Step-0 style) writes a ≤30-line report comparing prediction to outcome, stored in the evidence store.
Benefit: turns the evidence store from a log into a curriculum; improves both LLM and human proposal quality over time.
Gap: Mazda's loop is single-shot per failure; no orchestrated propose → experiment → gate cycling.
Integration: a run_improvement_loop.py entry point modeled on meta_harness.py (resume from summary, interrupt handling, per-iteration timing) — but its terminal state is “gated proposals awaiting human approval,” never activation. Budget caps (max iterations, max $ per run) enforced by CostGate.
Risks: cost runaway (cap it); proposal spam into the approval queue (dedupe via P2 history).
claude_wrapper.py; write a ~1-page mazda-proposer/SKILL.md (modeled on the text-classification SKILL.md: read history → hypothesize → emit structured JSON proposal; anti-leakage rules naming the holdout denylist; “one mechanism per candidate”).ImprovementProposal JSONs. Parse them into Mazda's Pydantic models; run each through BaselineCandidateExperimentRunner on the search fixtures and the existing GateChain, plus a prototype LeakageGate.MazdaProposalGenerator output for the same traces: gate-pass rate, val delta, human-rated proposal quality, and holdout delta for any gate-passer.experience/ and an evolution_summary.jsonl.propose_claude does) plus cheap deterministic evals; ~2–3 days of work.Do not replace Mazda with Meta-Harness — Meta-Harness itself says it's untested paper code, and it lacks every production control Mazda's invariants require (gates, approval, rollback, audit). Equally, do not leave Mazda's Propose station as a six-entry lookup table: it caps the system's ceiling at whatever failures its authors pre-imagined, and the missing holdout split means even those gains can't be trusted.
The right synthesis is Meta-Harness as the engine, Mazda as the brakes: adopt the val/test
firewall first (P0 — it's a correctness bug in Mazda today, not an enhancement), then swap in an LLM
proposer behind the existing IProposalGenerator port with a leakage gate, session logging, and
a search-history summary (P1–P4). Every adopted component enters through an existing port, is
selected in the composition root by profile, keeps the kernel finance-blind, keeps verification
deterministic, and ends at the human approval gateway — so all seven of Mazda's stated principles
survive intact.