Self-Improvement Systems · Comparative Review

Mazda vs. Meta-Harness

A review of Mazda's production self-improvement control plane against the Stanford IRIS Lab Meta-Harness framework (arXiv:2603.28052, github.com/stanford-iris-lab/meta-harness), with concrete, gate-preserving adoption recommendations.

No code modified during review Sources: mazda_dev_status.html + live Mazda source Sources: Meta-Harness repo (README, ONBOARDING, skills, loops)

Contents

1. Executive Summary

Mazda and Meta-Harness solve two different halves of the same problem, and each is strong exactly where the other is weak.

Mazda is a production control plane: a ports-and-adapters framework (nine factory families, six-station loop) that runs an agent, judges it deterministically, proposes a wrapper edit, A/B tests it, gates it (safety / cost / regression / usefulness), and activates it behind human approval with snapshots and rollback. Its governance is far ahead of Meta-Harness. But its generative half is thin: the live MazdaProposalGenerator is a hand-written lookup table mapping six FailureType values to six canned English patch strings. Mazda can only ever “propose” what a developer pre-wrote into that table. It also has no validation/test split, no cross-iteration search history, and no autonomous multi-iteration loop.

Meta-Harness is a research search engine: an outer loop where a full coding agent (Claude Code via claude_wrapper.py) reads scores and raw execution traces, forms falsifiable hypotheses, prototypes mechanisms, and writes arbitrary new harness code; an inner loop scores each candidate on a search set; a Pareto frontier and append-only evolution_summary.jsonl accumulate evidence across iterations; and a hard firewall separates validation (used during evolution) from held-out test (touched only once, at an explicit, irreversible finalization step). It has essentially no production governance: no gates beyond score comparison, no human approval, no activation/rollback, and it warns that candidates cost real money (~$500/iteration for TB2).

Recommendation. Keep Mazda's pipeline exactly as it is from Experiment onward, and replace/augment its Propose station and evidence-analysis loop with Meta-Harness's proposer pattern — an LLM proposer that emits ProposalCommand objects into Mazda's existing gate chain, plus a validation/test split, a frontier + evolution summary, and anti-leakage source checks. Nothing downstream of Propose needs to change, which is precisely what Mazda's interface-first design was built for.

2. Side-by-Side Comparison

DimensionMazda (agent_self_improvement)Meta-Harness
What it improves The wrapper around a fixed LLM: prompts, tool descriptions, context strategy, memory notes, workflows — for one live finance fleet The harness: arbitrary Python code around a fixed model (memory systems, agent scaffolds) — per domain, offline
Candidate generation Deterministic: FailureType → pre-written patch text (proposals.py); FailurePatternSuggester groups repeats (≥2) into one proposal LLM proposer agent (Claude Code, max effort) reads history + traces, forms 1–3 falsifiable hypotheses per iteration, mandatorily prototypes, writes new code
Search space Fixed menu of 4 command types over existing artifacts Arbitrary code; skill prompts forbid parameter-only variants and demand new mechanisms
Use of traces Traces persisted to SQLite; consumed by verifiers and pattern counting; not deeply read by any proposer Proposer deep-reads raw trajectories / log.jsonl as the primary analysis step; proposer's own sessions logged to experience/ (tools, files read, tokens, cost)
Historical evidence Evidence store holds traces / verdicts / proposals / approvals — audit-oriented, per-event evolution_summary.jsonl (per-candidate hypothesis, axis, delta, outcome, timing), frontier_val.json (best per dataset/task), per-iteration post-mortem reports — search-oriented, cross-iteration
Evaluation Deterministic-rules-first verifiers → ScoreCard → verdict; LLM judge only for ambiguity Task metric (accuracy / pass rate) on a search set; noise handled by multi-trial evals
Held-out test isolation Absent. A/B runs on task_inputs with no split; nothing prevents optimizing against the same cases forever Core design. Evolution sees only val; test results live in a directory never exposed; --test finalization freezes the run permanently (finalized.json blocks further evolution)
Overfitting prevention Only implicit (RegressionGate on baseline cases) Val/test split, anti-overfitting skill rules (no task names in code), universal + per-task forbidden_references denylist checked against candidate source before spending eval budget
Gates 4-gate chain: Safety, Cost, Regression, Usefulness — explicit contracts None as such; frontier selection is the only “gate”; TB2 loop does track cost/token rollout_metrics per candidate
Human approval / activation / rollback Approval gateway, snapshot store, activation + rollback services, audit records None. Winner = frontier top; adoption is manual / out-of-band
Trust boundary Layering rules (contracts import nothing concrete); generated code never exists — proposals are text patches Generated code is arbitrary Python, never imported by the controller; run in a short-lived sandboxed subprocess (Harbor pilot)
Autonomy One proposal per failing trace; no autonomous multi-iteration loop shipped Fully autonomous N-iteration loop with resume, fresh-start, interrupt handling
Cost accounting CostGate, per-run token usage in traces, estimated_cost_usd in experiments Per-candidate propose/bench/wall timing, tokens, $cost_usd per proposer session; explicit budget discipline (bring-up ladder: 1 task → 30 → 89)
Reproducibility SQLite evidence store, wrapper revisions by id, 200+ provider-free tests Run-name-isolated log dirs, full raw event logs, import-validation before eval, provider-free test suites
Maturity 4 of 9 factory families real; LLM/prompt/tool/context/workflow are stubs; loop not yet closed end-to-end in production “Cleaned-up paper code, not tested beyond verifying it runs”; research-grade

3. Reusable Meta-Harness Ideas, Classified

Adopt directly

  1. Validation / held-out test split with one-shot finalization (meta_harness.py: FINALIZED state machine, test results in a directory the proposer never sees).
  2. Append-only search history: evolution_summary.jsonl — one JSON row per candidate with hypothesis, axis, delta vs. frontier, outcome, timing — plus frontier_val.json.
  3. Proposer session logging (claude_wrapper.py's experience/ layout: meta.json, response.md, events.jsonl, per-tool-call files). Drop-in; it's a standalone file.
  4. Cheap pre-flight candidate validation before spending eval budget (import check / interface-compliance check, as in validate_candidates and TB2's validate_agent_class).

Adapt to Mazda

  1. LLM proposer agent replacing the lookup-table proposer — the single highest-value idea. An LlmProposalGenerator adapter behind the existing IProposalGenerator port, driven by a claude_wrapper-style subprocess with a Mazda-specific SKILL.md prior, emitting the same ImprovementProposal/command objects. Everything downstream (experiment, gates, approval, rollback) is untouched.
  2. Falsifiable hypothesis + prediction fields on every proposal, and “one mechanism per candidate” discipline (from both SKILL.md files). Maps onto Rationale.expected_benefit, which exists but is canned text today.
  3. Anti-leakage forbidden_references source check (Harbor controller.py: validate_source + universal denylist). Adapt as a fifth gate or a SafetyGate extension: proposed patch text must not reference held-out case identifiers, specific vendor names from eval fixtures, verifier internals, etc.
  4. Per-iteration post-mortem reports (Step 0 of the skill: ≤30-line “what changed, what improved/regressed, takeaway”). Adapt as an artifact the proposer writes into Mazda's evidence store/reporting family.
  5. Proposer prior as a versioned SKILL.md — the prior itself becomes part of the wrapper: versionable, diffable, and improvable through the same gates.
  6. Cost bring-up ladder (1 case → subset → full suite) and per-candidate rollout_metrics. Mazda's CostGate exists; the ladder is an evaluation-ordering policy to bolt onto the experiment runner.

Already present in Mazda

  1. Baseline-vs-candidate paired comparison with regression detection — BaselineCandidateExperimentRunner + ScorecardRegressionDetector are equivalent to (and cleaner than) Meta-Harness's frontier delta.
  2. Deterministic-rules-first evaluation — Mazda's Principle 5.1 is stronger than Meta-Harness's metric-only scoring.
  3. Never touch the base model — both systems share this thesis (harness ≡ wrapper).
  4. Sandbox/trust boundary for untrusted code — Mazda sidesteps it entirely by proposing text patches, not code; the Harbor subprocess isolation is only needed if Mazda ever lets a proposer write executable artifacts.
  5. Cost tracking primitives — token usage in traces, estimated_cost_usd, CostGate.

Not appropriate for Mazda

  1. Fully autonomous activation of frontier winners — violates Mazda's human-approval invariant. The autonomous loop should end at “proposal awaiting approval,” never at activation.
  2. Arbitrary-Python search space — Mazda's command-object granularity (prompt patch, tool-description patch, memory note, context rule) is the right blast radius for a live finance fleet. Letting the proposer write arbitrary code would demolish the gate chain's ability to reason about changes.
  3. “MUST produce N candidates every iteration, never stop” — right for research search under a paid budget, wrong for production: Mazda should be allowed to conclude “no material failure pattern this week; no proposal.”
  4. --fresh wipe-and-restart semantics — conflicts with Mazda's append-only audit evidence.

4. Prioritized Mazda Adoption Plan

Each item: idea → gap → integration point → benefit / risks → required evidence → priority & effort.

P0 — Validation/test split + finalization freeze

adopt · small

Gap: Mazda has no held-out isolation; repeated gate-passing against the same cases will silently overfit the wrapper to its own regression suite.

Integration: Split the evaluation case corpus (receipt fixtures, statement cases) into search and holdout sets at the persistence layer; ABRunConfig.task_inputs draws only from search. Add a FinalizationService in the improvement family: holdout is evaluated only when a candidate has passed all four gates and is queued for human approval, and a finalized flag on the wrapper-revision record blocks further tuning against that holdout snapshot.

Benefit: The UsefulnessGate's answer becomes trustworthy; approval decisions rest on unseen data.

Risks: Mazda's case corpus is small — a split shrinks statistical power. Mitigate with periodic holdout rotation (rotate, then retire the old holdout into search).

Evidence needed
Unit tests that evolution paths cannot read holdout results (mirror Meta-Harness's separate results/ dir); a demonstration that at least one past “improvement” scores differently on the split.
Effort
~2–4 days
Priority
Highest — correctness precondition for everything else

P1 — LLM proposer behind IProposalGenerator

adapt · medium

Gap: the deterministic proposer can only re-emit six pre-written sentences; it cannot learn from traces or invent a fix for an unanticipated failure (everything else falls into the MEMORY_NOTE catch-all).

Integration: new adapter implementations/improvement/llm_proposals.py::ClaudeProposalGenerator(IProposalGenerator) that (a) pulls the failing trace cluster + recent verdicts + current wrapper artifacts from the evidence store, (b) invokes a vendored claude_wrapper.run() with a Mazda SKILL.md prior, (c) parses a structured pending_proposal.json into the existing ImprovementProposal/ChangeSet command objects, (d) falls back to MazdaProposalGenerator on parse failure. Selected in the composition root via profile.improvement == "llm_proposer" — a profile edit, not a kernel change.

Benefit: proposals become trace-grounded and open-ended while remaining inspectable text patches that flow through the unchanged experiment/gate/approval/rollback pipeline. This is Meta-Harness's engine inside Mazda's brakes.

Risks: non-determinism in proposal content (acceptable: verdicts stay deterministic; the proposal is gated); prompt-injection via trace content into the proposer (constrain proposer tools to Read-only over an exported evidence snapshot, never live systems); token cost (log per-session cost à la claude_wrapper, enforce via CostGate).

Evidence needed
Offline replay — feed N historical failing traces to both proposers, human-rate proposal quality; verify every LLM proposal round-trips through Pydantic validation and the gate chain; verify fallback path.
Effort
~1–2 weeks
Priority
High — largest capability unlock

P2 — Search history + frontier

adopt · small

Gap: the evidence store records events but nothing summarizes what has been tried and with what delta, so a proposer (human or LLM) can't avoid re-proposing failed ideas.

Integration: an evolution_summary table (or JSONL exported by the reporting family) with per-proposal hypothesis, target artifact, gate outcomes, val delta; a frontier view of best-known wrapper revision per task type. Feed both into the P1 proposer's context.

Benefit: cross-iteration learning; the documented Meta-Harness failure mode (“parameter variants regress or tie”) is avoided by making history visible.

Evidence needed
Schema tests; proposer prompt includes and demonstrably references it.
Effort
2–3 days
Priority
High, immediately after P1 (they compound)

P3 — Anti-leakage source/patch check as a gate

adapt · small

Gap: nothing stops a proposal (especially an LLM one) from embedding holdout-case specifics (“Walgreens receipts always…”) — overfitting by memorization.

Integration: LeakageGate in implementations/evaluation/gates.py, configured with a denylist derived from holdout fixtures (vendor names, filenames, case ids) + a universal list, string-checked against patch_text — a direct port of controller.py::validate_source. Add it to the GateChain (making it five gates).

Benefit: cheap, deterministic, runs before any eval spend.

Risks: false positives on legitimately common substrings; keep the list reviewable in config.

Evidence needed
Unit tests with seeded leaks; run against all historical proposals (should pass).
Effort
1–2 days
Priority
Medium-high — mandatory before P1 goes live, trivial after P0 defines the holdout

P4 — Proposer session logging (experience/)

adopt · trivial

Gap: if P1 ships, the proposer's own reasoning process is unaudited — contrary to Mazda's auditability principle.

Integration: vendor claude_wrapper.py as-is; point log_dir at an experience/ sibling of the SQLite store; store the session directory path on the ImprovementProposal record.

Benefit: every activated change traces back to the exact proposer session (tools used, files read, tokens, cost) — extends Mazda's audit chain upstream into generation.

Effort
<1 day
Priority
Medium (bundled with P1)

P5 — Post-iteration reports + hypothesis/prediction fields

adapt · small

Gap: experiment outcomes aren't distilled into takeaways; expected_benefit is boilerplate, never scored against reality.

Integration: add hypothesis and prediction to Rationale; after each experiment, the reporting family (or the P1 proposer's next session, Step-0 style) writes a ≤30-line report comparing prediction to outcome, stored in the evidence store.

Benefit: turns the evidence store from a log into a curriculum; improves both LLM and human proposal quality over time.

Effort
2–3 days
Priority
Medium

P6 — Autonomous multi-iteration loop, ending at approval queue

adapt · medium

Gap: Mazda's loop is single-shot per failure; no orchestrated propose → experiment → gate cycling.

Integration: a run_improvement_loop.py entry point modeled on meta_harness.py (resume from summary, interrupt handling, per-iteration timing) — but its terminal state is “gated proposals awaiting human approval,” never activation. Budget caps (max iterations, max $ per run) enforced by CostGate.

Risks: cost runaway (cap it); proposal spam into the approval queue (dedupe via P2 history).

Evidence needed
Dry-run mode against the stub LLM family first; then a bounded live run.
Effort
~1 week
Priority
Lower — only worthwhile after P0–P3 exist

5. Proposed Low-Cost First Experiment

Offline proposer bake-off — zero risk to the live fleet, no activation, no fleet changes.
  1. Export 10–20 historical failing traces + verdicts from the SQLite evidence store into a read-only snapshot directory (this is the whole “world” the proposer may see).
  2. Split Mazda's existing evaluation fixtures ~70/30 into search/holdout (P0 in miniature, done by hand for the experiment).
  3. Vendor claude_wrapper.py; write a ~1-page mazda-proposer/SKILL.md (modeled on the text-classification SKILL.md: read history → hypothesize → emit structured JSON proposal; anti-leakage rules naming the holdout denylist; “one mechanism per candidate”).
  4. Run 3 iterations: each produces 2 candidate ImprovementProposal JSONs. Parse them into Mazda's Pydantic models; run each through BaselineCandidateExperimentRunner on the search fixtures and the existing GateChain, plus a prototype LeakageGate.
  5. Compare against the deterministic MazdaProposalGenerator output for the same traces: gate-pass rate, val delta, human-rated proposal quality, and holdout delta for any gate-passer.
  6. Log everything to experience/ and an evolution_summary.jsonl.
Cost
A few dollars of proposer tokens (or $0 on subscription auth, as propose_claude does) plus cheap deterministic evals; ~2–3 days of work.
Success criterion
At least one LLM-generated proposal passes all gates with a positive holdout delta that the lookup-table proposer could not have produced. That single artifact justifies (or kills) P1 with evidence instead of opinion.

6. Final Recommendation

Do not replace Mazda with Meta-Harness — Meta-Harness itself says it's untested paper code, and it lacks every production control Mazda's invariants require (gates, approval, rollback, audit). Equally, do not leave Mazda's Propose station as a six-entry lookup table: it caps the system's ceiling at whatever failures its authors pre-imagined, and the missing holdout split means even those gains can't be trusted.

The right synthesis is Meta-Harness as the engine, Mazda as the brakes: adopt the val/test firewall first (P0 — it's a correctness bug in Mazda today, not an enhancement), then swap in an LLM proposer behind the existing IProposalGenerator port with a leakage gate, session logging, and a search-history summary (P1–P4). Every adopted component enters through an existing port, is selected in the composition root by profile, keeps the kernel finance-blind, keeps verification deterministic, and ends at the human approval gateway — so all seven of Mazda's stated principles survive intact.

Classification for planning: P0, P2, P3, P4 are production-ready adoptions (small, deterministic, testable). P1, P5, P6 are staged experiments that must earn their way in through the Section 5 bake-off and Mazda's own gates.