⚠ SUPERSEDED (2026-06-15) — This "Mazda Orchestrator / Mazda gets direct finance tools" design is discarded. The current direction is that Mazda is the orchestrator, with minions that drive the Claude Agent SDK. See the canonical Mazda Dev Status doc (notes_plans_handoffs/mazda_dev_status.html). This page is retained for design history only and has been removed from the dashboard's Project Plans tab.
Project Plans - Mazda Orchestrator - GoF / SOLID / TDD

Mazda Orchestrator Team Construction Plan

Move Mazda from overloaded tool user to self-improving finance orchestrator. Mazda coordinates specialist agents, deterministic services, shared JSON contracts, evidence capture, evaluation gates, and rollback-safe improvement proposals.

Mazda Stays Self-Improving Orchestrator + Specialists Deterministic Services Stay Deterministic Abstract Factories Unit Tests First

1. Decision

Recommendation: Mazda becomes the manager/orchestrator agent for finance document runs. Specialist agents handle ambiguous, context-heavy work. Deterministic operations such as exact duplicate lookup, vendor-map lookup, trace writes, proposal persistence, DB insert, activation, and rollback remain service ports behind agents.

What Changes

  • Mazda stops carrying every operation as one overloaded personal toolset.
  • Mazda delegates to named stage agents through stable handoff contracts.
  • Every stage writes an artifact that can be replayed, tested, audited, or resumed.
  • Self-improvement is promoted from "extra tools" to a governed workflow lane.

What Does Not Change

  • Finance checks that must be exact stay in Python/SQL/YAML-backed services.
  • Only the storage lane writes final expense rows.
  • Improvement proposals never auto-activate without gates.
  • Mazda still records traces, evaluates outcomes, proposes changes, and learns.

2. Relationship To Existing Plans

This plan merges the team-agent idea from /home/adamsl/rol_finances/plans/team_construction_plan.md with the working pipeline details in /home/adamsl/rol_finances/plans/document_processing_steps.md.

It also uses the self-improvement architecture in Agent Self-Improvement System as the control plane. The most relevant pieces are:

Self-Improvement PhaseHow This Plan Uses It
Phase 02 - Factory ContractsAgent teams, services, and workflow objects are created from factories, not hard-coded into Mazda.
Phase 06 - Tool Contract LayerDeterministic services expose strict ports; agents do not invent or bypass tool contracts.
Phase 08 - Workflow SkeletonsThe document pipeline becomes a versioned workflow template with ordered steps and branch policy.
Phase 09/10 - Evaluation + GatesEvery run and every candidate improvement is scored, gated, and explained before activation.
Phase 11/12 - Proposals + ExperimentsImprovement Analyst drafts proposals; experiments compare baseline vs candidate workflows/prompts/tools.
Phase 13 - Activation + RollbackApproved changes activate through snapshots and rollback services, never by direct LLM mutation.
Phase 14 - Memory + NotesAccepted lessons update vendor/category/parser memory with source, reason, and staleness policy.
Phase 15/16 - Plugin Registry + ReportsNew agents and stage implementations register as plugins; reports show evidence, decisions, and active versions.

3. Target Architecture

MazdaOrchestratorAgent
  |
  +-- AgentTeamFactory
  |     +-- IntakeScanAgent
  |     +-- RouterAgent
  |     +-- ParserAgent
  |     +-- IdentityVendorAgent
  |     +-- DuplicateReviewAgent
  |     +-- ReceiptLinkerAgent
  |     +-- CategorizationAgent
  |     +-- StorageIngestAgent
  |     +-- ReviewExceptionAgent
  |     +-- ImprovementAnalystAgent
  |
  +-- DeterministicServiceFamily
        +-- TraceCommandService
        +-- VendorMapService
        +-- DuplicateLookupService
        +-- CategoryStoreService
        +-- ReceiptFinderService
        +-- ExpenseRepository
        +-- ProposalRepository
        +-- GateChain
        +-- ActivationRollbackService

Agents

Handle judgment, ambiguity, summarization, routing, exception packets, and explanation.

FacadeStrategyMediator

Services

Handle exact, auditable, repeatable operations: reads, writes, lookups, IDs, traces, gates.

CommandRepositoryUnit of Work

Tests

Pin every contract before implementation. Fakes prove orchestration without live Letta or live DB.

Contract TestsGolden CasesRegression Gates

4. Team Roles

RoleAgent / ServiceResponsibilitiesMust Not DoPrimary Patterns
Mazda Orchestrator Agent Own run state, delegate stages, enforce workflow order, collect artifacts, request review, ask for improvement analysis. Parse documents directly, write final DB rows, activate self-changes directly. Facade, Mediator, Template Method
Intake / Scan Agent + file service Create run folder, manifest, previews, source metadata. Guess categories or mutate finance DB. Factory Method, Adapter
Router Agent Classify document type and select parser strategy. Extract final rows or write storage. Strategy, Chain of Responsibility
Parser Agent + parser adapters Use OCR, Gemini, or structured parsers to emit transaction candidates with evidence. Invent missing fields or insert rows. Adapter, Strategy
Identity + Vendor Agent + deterministic vendor service Normalize vendor before final id_light; explain alias choice. Bypass vendor map or create unstable IDs. Strategy, Specification
Duplicate Review Agent + deterministic duplicate service Interpret exact/fuzzy duplicate results and recommend skip/review/proceed. Perform approximate DB logic inside the LLM. Chain of Responsibility, Policy Object
Receipt Linker Agent + receipt finder service Link best receipt/document artifact; surface alternatives. Write final rows. Strategy, Adapter
Categorization Agent + category services Use vendor store first, categorizer second, review third; learn from accepted decisions. Silently insert null or low-confidence categories. Chain of Responsibility, Strategy
Storage / Ingest Agent gate + repository service Validate required fields, write final DB/filesystem records, emit ingest summary. Accept ambiguous rows without review. Unit of Work, Repository, Command
Review / Exception Agent Prepare human review packets and turn accepted corrections into lessons. Auto-approve risky corrections. State, Observer
Trace Steward Service first, optional agent facade Validate trace packet shape, then call deterministic record_trace. Let an LLM invent audit rows. Command, Memento, Repository
Improvement Analyst Agent + proposal service Analyze failed traces and draft proposal candidates with risk/tests/benefit. Persist or activate proposals without the proposal service and gates. Builder, Command, Chain of Responsibility

5. Canonical Workflow

intake
  -> route
  -> parse
  -> normalize_vendor_and_identity
  -> duplicate_check
  -> receipt_link
  -> categorize
  -> storage_validation
  -> store
  -> record_trace
  -> evaluate
  -> propose_improvement_if_reproducible
  -> experiment_and_gate
  -> activate_or_reject
  -> report_and_learn

The workflow should be represented as IWorkflow / IStep objects, not as scattered prose in tool descriptions. Mazda can read the workflow revision and run it as a template. Steps can be swapped by document type while the skeleton stays stable.

6. Interface Catalog

Orchestration

interface IOrchestratorAgent {
  runDocumentJob(request: DocumentJobRequest): OrchestratorRunResult
}

interface IAgentTeamFactory {
  createTeam(profile: TeamRuntimeProfile): AgentTeam
}

interface IWorkflowRunner {
  run(workflow: IWorkflow, context: WorkflowContext): WorkflowResult
}

Agent Ports

interface IStageAgent<TIn, TOut> {
  stageName(): string
  run(input: StageEnvelope<TIn>): StageEnvelope<TOut>
}

interface IAgentTransport {
  send(agentId: AgentId, request: AgentRunRequest): AgentRunResult
}

Artifacts

interface IRunArtifactStore {
  put(runId: RunId, artifact: Artifact): ArtifactRef
  get(ref: ArtifactRef): Artifact
  list(runId: RunId): ArtifactRef[]
}

interface IStageEnvelopeValidator {
  validate(envelope: StageEnvelope<unknown>): ValidationResult
}

Deterministic Finance Services

interface IVendorMapService {
  resolve(rawDescription: string): VendorResolution
}

interface IDuplicateLookupService {
  check(candidate: CanonicalTransaction): DuplicateVerdict
}

interface ICategoryResolutionService {
  resolve(candidate: CanonicalTransaction): CategoryVerdict
}

Storage

interface IExpenseUnitOfWork {
  validate(batch: ValidatedExpenseBatch): ValidationResult
  commit(batch: ValidatedExpenseBatch): IngestResult
  rollback(snapshot: SnapshotId): RollbackResult
}

interface IStoragePolicy {
  canStore(candidate: CanonicalTransaction): PolicyVerdict
}

Self-Improvement

interface ITraceCommandService {
  record(packet: TracePacket): TraceId
}

interface IImprovementAnalystAgent {
  analyze(trace: TraceRecord): ImprovementDraft
}

interface IProposalCommandService {
  file(draft: ImprovementDraft): ProposalId
}

Evaluation + Gates

interface IEvaluationSuite {
  score(run: OrchestratorRunResult): ScoreCard
}

interface IGate {
  evaluate(candidate: ImprovementCandidate): GateVerdict
}

interface IGateChain {
  evaluate(candidate: ImprovementCandidate): FinalVerdict
}

Reporting + Learning

interface IRunReportBuilder {
  build(run: OrchestratorRunResult): HtmlReport
}

interface ILessonWriter {
  write(lesson: AcceptedLesson): LessonId
}

7. Shared JSON Envelope

Every agent handoff uses one envelope so Mazda can retry, audit, replay, and resume without understanding each specialist's internals.

{
  "schema_version": "finance-stage-envelope/v1",
  "run_id": "run_...",
  "document_id": "doc_...",
  "stage": "parse",
  "status": "ok|needs_review|failed",
  "confidence": 0.97,
  "warnings": [],
  "artifact_refs": [],
  "data": {},
  "evidence": [],
  "next_recommended_stage": "normalize_vendor_and_identity"
}

8. GoF Pattern Map

PatternWhere It GoesReason
FacadeMazdaOrchestratorAgentExpose one simple run API while hiding team complexity.
MediatorMazda coordinating agentsSpecialists do not directly depend on each other.
Abstract FactoryIAgentTeamFactory, IServiceFamilyFactoryCreate compatible agent/service families by runtime profile.
Template MethodDocumentWorkflowTemplateStable finance workflow with overridable stage behavior.
StrategyRouting, parsing, category resolution, receipt linkingSwap algorithms by document type and confidence.
AdapterLetta agents, Gemini parser, OCR, existing finance modulesNormalize external APIs into internal ports.
CommandTrace write, proposal file, DB insert, activation, rollbackAuditable, replayable actions.
Chain of ResponsibilityCategory resolution, duplicate policy, improvement gatesFast deterministic checks first, expensive/risky checks later.
RepositoryArtifacts, traces, proposals, snapshots, expensesPersistence stays behind contracts.
Unit of WorkStorage ingest and activationCommit/rollback entire validated changesets.
MementoWorkflow/prompt/tool snapshots before activationRollback to previous active versions.
ObserverRun events and dashboard statusReports and status update without coupling to execution.
StateDocument job lifecycleExplicit transitions: queued, running, review, stored, failed, improved.

9. Unit Test Plan

TDD rule: Start with fake agents, fake transports, temp artifact stores, and in-memory repositories. Live Letta and live finance DB tests come later as integration tests.
Test FileBehavior To PinFirst Failing Assertion
test_orchestrator_workflow_order.pyMazda calls stages in canonical order.Step sequence equals intake, route, parse, normalize, duplicate, receipt, categorize, store, trace, evaluate.
test_agent_team_factory.pyFactory builds a complete team for the Mazda profile.All required stage agents and services are non-null and compatible.
test_stage_envelope_contract.pyEvery stage returns the shared JSON envelope.Missing run_id, stage, status, or confidence fails validation.
test_parser_stage_contract.pyParser emits candidates with evidence and warnings.Parser may not invent required fields without evidence.
test_vendor_identity_order.pyFinal id_light is built after canonical vendor resolution.id_light.vendor equals canonical vendor key, not raw description guess.
test_duplicate_policy.pyExact duplicates stop storage; fuzzy duplicates route to review unless high-confidence safe.Exact duplicate never reaches storage stage.
test_category_chain.pyVendor store first, categorizer second, review third.Known vendor does not call LLM categorizer.
test_storage_single_writer.pyOnly storage/ingest service can commit final expense rows.Parser/router/category agents have no DB write port.
test_trace_command_once.pyTrace is recorded exactly once per orchestrator run.Success and failure runs each produce one trace packet.
test_improvement_gate.pyImprovement proposal requires reproducible failure and trace id.Clean run does not call proposal service.
test_activation_rollback.pyApproved changes snapshot before activation and can roll back.Activation without snapshot fails.
test_memory_lesson_policy.pyAccepted lessons include source, reason, owner, and staleness policy.Anonymous memory update fails.

10. Example Unit Tests

Workflow Order

def test_mazda_runs_canonical_workflow_order():
    team = FakeAgentTeam()
    runner = MazdaOrchestrator(team=team, services=FakeServices())

    runner.run_document_job(sample_statement_job())

    assert team.called_stages == [
        "intake", "route", "parse", "normalize_vendor_and_identity",
        "duplicate_check", "receipt_link", "categorize",
        "storage_validation", "store"
    ]

No DB Writes Outside Storage

def test_only_storage_stage_receives_expense_repository():
    factory = MazdaTeamFactory(profile="test")

    team = factory.create_team()

    assert not hasattr(team.parser_agent, "expense_repository")
    assert not hasattr(team.router_agent, "expense_repository")
    assert team.storage_agent.expense_repository is not None

Self-Improvement Gate

def test_clean_run_does_not_file_improvement():
    services = FakeServices()
    runner = MazdaOrchestrator(team=CleanRunTeam(), services=services)

    runner.run_document_job(sample_statement_job())

    assert services.trace_command.calls == 1
    assert services.proposal_command.calls == 0

Proposal Requires Trace

def test_improvement_draft_requires_trace_id():
    draft = ImprovementDraft(trace_id=None, failure_type="wrong_category")

    result = ProposalPolicy().validate(draft)

    assert result.ok is False
    assert "trace_id" in result.reason

11. Build Phases

PhaseGoalInterfaces / ClassesTests FirstDone Means
00
Vocabulary
Define IDs, envelopes, artifact refs, stage names. RunId, DocumentId, StageName, StageEnvelope test_stage_envelope_contract.py All handoffs share one schema.
01
Factories
Create team and service families. IAgentTeamFactory, IServiceFamilyFactory, TeamRuntimeProfile test_agent_team_factory.py Mazda receives complete team through interfaces.
02
Workflow Skeleton
Represent canonical finance workflow as steps. IWorkflow, IStep, IWorkflowRunner test_orchestrator_workflow_order.py Stage order is testable and versioned.
03
Fake Agent Runtime
Run full workflow offline with fake agents. FakeStageAgent, FakeAgentTransport, FakeArtifactStore Golden fake-run test No live Letta dependency for unit tests.
04
Deterministic Ports
Wrap existing finance modules behind services. IVendorMapService, IDuplicateLookupService, IExpenseUnitOfWork Vendor, duplicate, category, storage contract tests Agents cannot bypass deterministic checks.
05
Self-Improvement Lane
Trace, evaluate, propose, gate, activate, rollback. ITraceCommandService, IImprovementAnalystAgent, IGateChain Trace-once and clean-run-no-proposal tests Mazda self-improves safely after evidence.
06
Live Letta Wiring
Bind stage agents to real Letta IDs through adapters. ILettaAgentAdapter, IAgentRegistry Registration and transport integration tests Specialists can be swapped without Mazda code changes.
07
Dashboard Reports
Show run state, artifacts, gates, active versions. IRunReportBuilder, IStatusObserver Snapshot-stable report tests Team can review outcomes without reading logs.

12. Open Design Decisions

Agent Granularity

Start with existing agents covering multiple roles. Split into new agents only when repeated workload, context size, or failure isolation justifies it.

Trace Steward Shape

Prefer deterministic record_trace. Add a Trace Steward Agent only to inspect/complete evidence packets before the command service writes.

Improvement Analyst Scope

The analyst may draft changes, but proposal persistence, gates, experiments, activation, and rollback remain deterministic services.

13. Immediate Action Items

  1. Approve Mazda as the orchestrator and the canonical step order.
  2. Create StageEnvelope, AgentTeam, and WorkflowContext contracts.
  3. Add unit tests for workflow order, envelope validation, single-writer storage, trace-once behavior, and clean-run no-proposal behavior.
  4. Wrap current deterministic finance functions as services before adding more live agents.
  5. Add fake specialist agents and prove a full offline run.
  6. Wire existing Letta agents through an agent registry and transport adapter.
  7. Add the self-improvement lane: trace, score, draft proposal, gate, experiment, activate/rollback.
  8. Publish run reports in the dashboard so the team can review evidence and active versions.