1. Decision
What Changes
- Mazda stops carrying every operation as one overloaded personal toolset.
- Mazda delegates to named stage agents through stable handoff contracts.
- Every stage writes an artifact that can be replayed, tested, audited, or resumed.
- Self-improvement is promoted from "extra tools" to a governed workflow lane.
What Does Not Change
- Finance checks that must be exact stay in Python/SQL/YAML-backed services.
- Only the storage lane writes final expense rows.
- Improvement proposals never auto-activate without gates.
- Mazda still records traces, evaluates outcomes, proposes changes, and learns.
2. Relationship To Existing Plans
This plan merges the team-agent idea from /home/adamsl/rol_finances/plans/team_construction_plan.md with the working pipeline details in /home/adamsl/rol_finances/plans/document_processing_steps.md.
It also uses the self-improvement architecture in Agent Self-Improvement System as the control plane. The most relevant pieces are:
| Self-Improvement Phase | How This Plan Uses It |
|---|---|
| Phase 02 - Factory Contracts | Agent teams, services, and workflow objects are created from factories, not hard-coded into Mazda. |
| Phase 06 - Tool Contract Layer | Deterministic services expose strict ports; agents do not invent or bypass tool contracts. |
| Phase 08 - Workflow Skeletons | The document pipeline becomes a versioned workflow template with ordered steps and branch policy. |
| Phase 09/10 - Evaluation + Gates | Every run and every candidate improvement is scored, gated, and explained before activation. |
| Phase 11/12 - Proposals + Experiments | Improvement Analyst drafts proposals; experiments compare baseline vs candidate workflows/prompts/tools. |
| Phase 13 - Activation + Rollback | Approved changes activate through snapshots and rollback services, never by direct LLM mutation. |
| Phase 14 - Memory + Notes | Accepted lessons update vendor/category/parser memory with source, reason, and staleness policy. |
| Phase 15/16 - Plugin Registry + Reports | New agents and stage implementations register as plugins; reports show evidence, decisions, and active versions. |
3. Target Architecture
MazdaOrchestratorAgent
|
+-- AgentTeamFactory
| +-- IntakeScanAgent
| +-- RouterAgent
| +-- ParserAgent
| +-- IdentityVendorAgent
| +-- DuplicateReviewAgent
| +-- ReceiptLinkerAgent
| +-- CategorizationAgent
| +-- StorageIngestAgent
| +-- ReviewExceptionAgent
| +-- ImprovementAnalystAgent
|
+-- DeterministicServiceFamily
+-- TraceCommandService
+-- VendorMapService
+-- DuplicateLookupService
+-- CategoryStoreService
+-- ReceiptFinderService
+-- ExpenseRepository
+-- ProposalRepository
+-- GateChain
+-- ActivationRollbackServiceAgents
Handle judgment, ambiguity, summarization, routing, exception packets, and explanation.
FacadeStrategyMediatorServices
Handle exact, auditable, repeatable operations: reads, writes, lookups, IDs, traces, gates.
CommandRepositoryUnit of WorkTests
Pin every contract before implementation. Fakes prove orchestration without live Letta or live DB.
Contract TestsGolden CasesRegression Gates4. Team Roles
| Role | Agent / Service | Responsibilities | Must Not Do | Primary Patterns |
|---|---|---|---|---|
| Mazda Orchestrator | Agent | Own run state, delegate stages, enforce workflow order, collect artifacts, request review, ask for improvement analysis. | Parse documents directly, write final DB rows, activate self-changes directly. | Facade, Mediator, Template Method |
| Intake / Scan | Agent + file service | Create run folder, manifest, previews, source metadata. | Guess categories or mutate finance DB. | Factory Method, Adapter |
| Router | Agent | Classify document type and select parser strategy. | Extract final rows or write storage. | Strategy, Chain of Responsibility |
| Parser | Agent + parser adapters | Use OCR, Gemini, or structured parsers to emit transaction candidates with evidence. | Invent missing fields or insert rows. | Adapter, Strategy |
| Identity + Vendor | Agent + deterministic vendor service | Normalize vendor before final id_light; explain alias choice. |
Bypass vendor map or create unstable IDs. | Strategy, Specification |
| Duplicate Review | Agent + deterministic duplicate service | Interpret exact/fuzzy duplicate results and recommend skip/review/proceed. | Perform approximate DB logic inside the LLM. | Chain of Responsibility, Policy Object |
| Receipt Linker | Agent + receipt finder service | Link best receipt/document artifact; surface alternatives. | Write final rows. | Strategy, Adapter |
| Categorization | Agent + category services | Use vendor store first, categorizer second, review third; learn from accepted decisions. | Silently insert null or low-confidence categories. | Chain of Responsibility, Strategy |
| Storage / Ingest | Agent gate + repository service | Validate required fields, write final DB/filesystem records, emit ingest summary. | Accept ambiguous rows without review. | Unit of Work, Repository, Command |
| Review / Exception | Agent | Prepare human review packets and turn accepted corrections into lessons. | Auto-approve risky corrections. | State, Observer |
| Trace Steward | Service first, optional agent facade | Validate trace packet shape, then call deterministic record_trace. |
Let an LLM invent audit rows. | Command, Memento, Repository |
| Improvement Analyst | Agent + proposal service | Analyze failed traces and draft proposal candidates with risk/tests/benefit. | Persist or activate proposals without the proposal service and gates. | Builder, Command, Chain of Responsibility |
5. Canonical Workflow
intake
-> route
-> parse
-> normalize_vendor_and_identity
-> duplicate_check
-> receipt_link
-> categorize
-> storage_validation
-> store
-> record_trace
-> evaluate
-> propose_improvement_if_reproducible
-> experiment_and_gate
-> activate_or_reject
-> report_and_learnThe workflow should be represented as IWorkflow / IStep objects, not as scattered prose in tool descriptions. Mazda can read the workflow revision and run it as a template. Steps can be swapped by document type while the skeleton stays stable.
6. Interface Catalog
Orchestration
interface IOrchestratorAgent {
runDocumentJob(request: DocumentJobRequest): OrchestratorRunResult
}
interface IAgentTeamFactory {
createTeam(profile: TeamRuntimeProfile): AgentTeam
}
interface IWorkflowRunner {
run(workflow: IWorkflow, context: WorkflowContext): WorkflowResult
}
Agent Ports
interface IStageAgent<TIn, TOut> {
stageName(): string
run(input: StageEnvelope<TIn>): StageEnvelope<TOut>
}
interface IAgentTransport {
send(agentId: AgentId, request: AgentRunRequest): AgentRunResult
}
Artifacts
interface IRunArtifactStore {
put(runId: RunId, artifact: Artifact): ArtifactRef
get(ref: ArtifactRef): Artifact
list(runId: RunId): ArtifactRef[]
}
interface IStageEnvelopeValidator {
validate(envelope: StageEnvelope<unknown>): ValidationResult
}
Deterministic Finance Services
interface IVendorMapService {
resolve(rawDescription: string): VendorResolution
}
interface IDuplicateLookupService {
check(candidate: CanonicalTransaction): DuplicateVerdict
}
interface ICategoryResolutionService {
resolve(candidate: CanonicalTransaction): CategoryVerdict
}
Storage
interface IExpenseUnitOfWork {
validate(batch: ValidatedExpenseBatch): ValidationResult
commit(batch: ValidatedExpenseBatch): IngestResult
rollback(snapshot: SnapshotId): RollbackResult
}
interface IStoragePolicy {
canStore(candidate: CanonicalTransaction): PolicyVerdict
}
Self-Improvement
interface ITraceCommandService {
record(packet: TracePacket): TraceId
}
interface IImprovementAnalystAgent {
analyze(trace: TraceRecord): ImprovementDraft
}
interface IProposalCommandService {
file(draft: ImprovementDraft): ProposalId
}
Evaluation + Gates
interface IEvaluationSuite {
score(run: OrchestratorRunResult): ScoreCard
}
interface IGate {
evaluate(candidate: ImprovementCandidate): GateVerdict
}
interface IGateChain {
evaluate(candidate: ImprovementCandidate): FinalVerdict
}
Reporting + Learning
interface IRunReportBuilder {
build(run: OrchestratorRunResult): HtmlReport
}
interface ILessonWriter {
write(lesson: AcceptedLesson): LessonId
}
7. Shared JSON Envelope
Every agent handoff uses one envelope so Mazda can retry, audit, replay, and resume without understanding each specialist's internals.
{
"schema_version": "finance-stage-envelope/v1",
"run_id": "run_...",
"document_id": "doc_...",
"stage": "parse",
"status": "ok|needs_review|failed",
"confidence": 0.97,
"warnings": [],
"artifact_refs": [],
"data": {},
"evidence": [],
"next_recommended_stage": "normalize_vendor_and_identity"
}
8. GoF Pattern Map
| Pattern | Where It Goes | Reason |
|---|---|---|
| Facade | MazdaOrchestratorAgent | Expose one simple run API while hiding team complexity. |
| Mediator | Mazda coordinating agents | Specialists do not directly depend on each other. |
| Abstract Factory | IAgentTeamFactory, IServiceFamilyFactory | Create compatible agent/service families by runtime profile. |
| Template Method | DocumentWorkflowTemplate | Stable finance workflow with overridable stage behavior. |
| Strategy | Routing, parsing, category resolution, receipt linking | Swap algorithms by document type and confidence. |
| Adapter | Letta agents, Gemini parser, OCR, existing finance modules | Normalize external APIs into internal ports. |
| Command | Trace write, proposal file, DB insert, activation, rollback | Auditable, replayable actions. |
| Chain of Responsibility | Category resolution, duplicate policy, improvement gates | Fast deterministic checks first, expensive/risky checks later. |
| Repository | Artifacts, traces, proposals, snapshots, expenses | Persistence stays behind contracts. |
| Unit of Work | Storage ingest and activation | Commit/rollback entire validated changesets. |
| Memento | Workflow/prompt/tool snapshots before activation | Rollback to previous active versions. |
| Observer | Run events and dashboard status | Reports and status update without coupling to execution. |
| State | Document job lifecycle | Explicit transitions: queued, running, review, stored, failed, improved. |
9. Unit Test Plan
| Test File | Behavior To Pin | First Failing Assertion |
|---|---|---|
test_orchestrator_workflow_order.py | Mazda calls stages in canonical order. | Step sequence equals intake, route, parse, normalize, duplicate, receipt, categorize, store, trace, evaluate. |
test_agent_team_factory.py | Factory builds a complete team for the Mazda profile. | All required stage agents and services are non-null and compatible. |
test_stage_envelope_contract.py | Every stage returns the shared JSON envelope. | Missing run_id, stage, status, or confidence fails validation. |
test_parser_stage_contract.py | Parser emits candidates with evidence and warnings. | Parser may not invent required fields without evidence. |
test_vendor_identity_order.py | Final id_light is built after canonical vendor resolution. | id_light.vendor equals canonical vendor key, not raw description guess. |
test_duplicate_policy.py | Exact duplicates stop storage; fuzzy duplicates route to review unless high-confidence safe. | Exact duplicate never reaches storage stage. |
test_category_chain.py | Vendor store first, categorizer second, review third. | Known vendor does not call LLM categorizer. |
test_storage_single_writer.py | Only storage/ingest service can commit final expense rows. | Parser/router/category agents have no DB write port. |
test_trace_command_once.py | Trace is recorded exactly once per orchestrator run. | Success and failure runs each produce one trace packet. |
test_improvement_gate.py | Improvement proposal requires reproducible failure and trace id. | Clean run does not call proposal service. |
test_activation_rollback.py | Approved changes snapshot before activation and can roll back. | Activation without snapshot fails. |
test_memory_lesson_policy.py | Accepted lessons include source, reason, owner, and staleness policy. | Anonymous memory update fails. |
10. Example Unit Tests
Workflow Order
def test_mazda_runs_canonical_workflow_order():
team = FakeAgentTeam()
runner = MazdaOrchestrator(team=team, services=FakeServices())
runner.run_document_job(sample_statement_job())
assert team.called_stages == [
"intake", "route", "parse", "normalize_vendor_and_identity",
"duplicate_check", "receipt_link", "categorize",
"storage_validation", "store"
]
No DB Writes Outside Storage
def test_only_storage_stage_receives_expense_repository():
factory = MazdaTeamFactory(profile="test")
team = factory.create_team()
assert not hasattr(team.parser_agent, "expense_repository")
assert not hasattr(team.router_agent, "expense_repository")
assert team.storage_agent.expense_repository is not None
Self-Improvement Gate
def test_clean_run_does_not_file_improvement():
services = FakeServices()
runner = MazdaOrchestrator(team=CleanRunTeam(), services=services)
runner.run_document_job(sample_statement_job())
assert services.trace_command.calls == 1
assert services.proposal_command.calls == 0
Proposal Requires Trace
def test_improvement_draft_requires_trace_id():
draft = ImprovementDraft(trace_id=None, failure_type="wrong_category")
result = ProposalPolicy().validate(draft)
assert result.ok is False
assert "trace_id" in result.reason
11. Build Phases
| Phase | Goal | Interfaces / Classes | Tests First | Done Means |
|---|---|---|---|---|
| 00 Vocabulary |
Define IDs, envelopes, artifact refs, stage names. | RunId, DocumentId, StageName, StageEnvelope |
test_stage_envelope_contract.py |
All handoffs share one schema. |
| 01 Factories |
Create team and service families. | IAgentTeamFactory, IServiceFamilyFactory, TeamRuntimeProfile |
test_agent_team_factory.py |
Mazda receives complete team through interfaces. |
| 02 Workflow Skeleton |
Represent canonical finance workflow as steps. | IWorkflow, IStep, IWorkflowRunner |
test_orchestrator_workflow_order.py |
Stage order is testable and versioned. |
| 03 Fake Agent Runtime |
Run full workflow offline with fake agents. | FakeStageAgent, FakeAgentTransport, FakeArtifactStore |
Golden fake-run test | No live Letta dependency for unit tests. |
| 04 Deterministic Ports |
Wrap existing finance modules behind services. | IVendorMapService, IDuplicateLookupService, IExpenseUnitOfWork |
Vendor, duplicate, category, storage contract tests | Agents cannot bypass deterministic checks. |
| 05 Self-Improvement Lane |
Trace, evaluate, propose, gate, activate, rollback. | ITraceCommandService, IImprovementAnalystAgent, IGateChain |
Trace-once and clean-run-no-proposal tests | Mazda self-improves safely after evidence. |
| 06 Live Letta Wiring |
Bind stage agents to real Letta IDs through adapters. | ILettaAgentAdapter, IAgentRegistry |
Registration and transport integration tests | Specialists can be swapped without Mazda code changes. |
| 07 Dashboard Reports |
Show run state, artifacts, gates, active versions. | IRunReportBuilder, IStatusObserver |
Snapshot-stable report tests | Team can review outcomes without reading logs. |
12. Open Design Decisions
Agent Granularity
Start with existing agents covering multiple roles. Split into new agents only when repeated workload, context size, or failure isolation justifies it.
Trace Steward Shape
Prefer deterministic record_trace. Add a Trace Steward Agent only to inspect/complete evidence packets before the command service writes.
Improvement Analyst Scope
The analyst may draft changes, but proposal persistence, gates, experiments, activation, and rollback remain deterministic services.
13. Immediate Action Items
- Approve Mazda as the orchestrator and the canonical step order.
- Create
StageEnvelope,AgentTeam, andWorkflowContextcontracts. - Add unit tests for workflow order, envelope validation, single-writer storage, trace-once behavior, and clean-run no-proposal behavior.
- Wrap current deterministic finance functions as services before adding more live agents.
- Add fake specialist agents and prove a full offline run.
- Wire existing Letta agents through an agent registry and transport adapter.
- Add the self-improvement lane: trace, score, draft proposal, gate, experiment, activate/rollback.
- Publish run reports in the dashboard so the team can review evidence and active versions.