Documentation of the primary data structures and objects Mazda uses during document intake and processing workflows.
Core Data Objects
IntakeVerificationEvidence
A structured JSON object that Mazda records at STEP 5 of the document intake pipeline to persist evidence for the self-improvement loop.
Purpose: Provides deterministic, verifiable records of what Mazda did, what tools returned, and what decisions were made.
Used by the judge (STEP 6) to assess whether the run succeeded or failed.
task_name: string — always "document-intake"
wrapper_revision: string — loaded via load_wrapper_revision; her active learned rules at this moment
tool_calls: array — every tool invocation with arguments and return values
step_results: object — outcome of each STEP (1–9), including success/failure and why
trace_id: integer — returned by record_trace; used as a key for rollback/proposal filing
Expense
The fundamental record in the ROL Finance database representing a single financial transaction.
id: integer — primary key in MySQL
expense_date: date — transaction date (YYYY-MM-DD)
amount: decimal — transaction amount (signed; negative = refund)
vendor_key: string — standardized vendor identifier (e.g., consumers_7996)
description: string — human-readable transaction text
category: string — expense category (e.g., utilities, meals); null if unknown
receipt_url: string — optional path to receipt file
Statement Transaction Row
A single transaction extracted from a bank statement, before deduplication and storage.
date: date — when the transaction posted
amount: decimal — transaction amount
description: string — bank's description (may contain merchant name, reference, etc.)
status: string — "readable" or "unknown" (if the amount couldn't be parsed)
DocumentIntakeResult
Returned by the deterministic text-extraction facade (mazda_intake.py) after analyzing a scanned document.
ok: boolean — true if extraction succeeded
doc_kind: string — document classification: statement, receipt, invoice, or unknown
confidence: float — 0.0–1.0; how certain the facade is about the classification
action: string — recommended next step: store, reject, manual_review
extracted_data: object (optional) — raw text or structured fields extracted from the document
vendor: string (optional) — detected vendor/merchant name (for receipts/invoices)
Processing Pipeline Objects
MazdalntakeFacade
The deterministic text-extraction strategy interface; abstracts away the choice of OCR, vision model, or direct PDF parsing.
Key invariant: Returns ok:true, doc_kind:unknown, confidence:0 for scanned JPEGs with no extractable text.
When doc_kind==unknown OR confidence==0, Mazda is routed to classify the image herself.
process(file_path): method → DocumentIntakeResult
supports_format(ext): method → boolean (checks if facade can handle JPG, PDF, PNG, etc.)
Document Classification Result
Returned when Mazda runs vision classification (Gemini) on a scanned image to determine doc_kind and vendor.
doc_type or doc_kind: string — vision model's best guess at document type
merchant or vendor: string — inferred vendor/merchant name
confidence: float — model's confidence in the classification
key_details: object (optional) — any structured fields the model extracted (amounts, dates, etc.)
Naming divergence: mazda_intake.py uses doc_kind/vendor, while
rol_finances/tools/classify_scan.py (Mazda's vision classifier) uses doc_type/merchant.
Merge functions accept either.
Verification Objects
Duplicate Check Result
Returned by check_duplicates after querying the database for existing matching rows.
matches: array of objects — rows in the DB matching on (expense_date, amount)
match[].id: integer — existing expense ID
match[].id_light: string — normalized key (e.g., consumers_energy_01_23_25_222_65)
match[].vendor_key: string — vendor prefix from the stored row
is_duplicate: boolean — true if an exact or fuzzy match exists
Vendor Validation Result
Returned by check_vendor_key after looking up a vendor in the standard vendor map.
recognized: boolean — true if vendor_key exists in the map
vendor_key: string — the normalized key
expected_category: string — what category this vendor typically uses
aliases: array (optional) — alternative names this vendor goes by
Category Validation Result
Returned by check_category after verifying that an expense category matches the vendor's expectations.
matches_vendor: boolean — true if the category is acceptable for this vendor
expected_category: string — what the vendor map says
provided_category: string — what was submitted
suggestion: string (optional) — if mismatch, a better category choice
Statement Totals Verification
Returned by verify_statement_totals after checking that individual transaction amounts sum to the reported total.
ok: boolean — true if sum of rows == reported total (within rounding)
reported_total: decimal — what the statement header claims
computed_total: decimal — sum of all transaction amounts
difference: decimal — reported − computed (should be ~0)
message: string — human-readable explanation of the result
Self-Improvement & Governance Objects
Wrapper Revision
Loaded at the start of a run via load_wrapper_revision; contains Mazda's accumulated learned rules and configuration.
Key invariant: Load this once at STEP 1 and keep the returned revision IDs for record_trace.
The instructions block is immutable; a new revision is created when a proposal is activated.
found: boolean — whether the agent/revision exists
wrapper_revision: string — unique ID for this revision snapshot
system_message_revision: string — ID of the system prompt this revision uses
toolset_revision: string — ID of the tool set
is_active: boolean — whether this is the currently-active revision
instructions: string — the accumulated LEARNED RULES; must be read and applied at every relevant step
Trace Record
Persisted via record_trace to the self-improvement SQLite database; contains the complete evidence of a run.
trace_id: integer — primary key; returned immediately after record
agent_name: string — always "Mazda"
evidence_json: JSON — the IntakeVerificationEvidence object serialized
verdict: string — PASS, FAIL, or NEEDS_REVIEW (set by the judge in STEP 6)
created_at: datetime — when the record was filed
Improvement Proposal
Filed via propose_improvement when a run fails; describes what went wrong and suggests a fix.
proposal_id: integer — unique identifier
trace_id: integer — links back to the failed run's evidence
agent_name: string — "Mazda"
problem_description: string — what went wrong
suggested_instruction: string — one imperative rule that would have prevented this
risk_level: string — low, medium, or high
status: string — PENDING, ACTIVATED, REJECTED, or PENDING_APPROVAL
Example instruction: "For doc_kind=statement, always run parse_statement_scan.py then store_statement_transactions.py"
4-Gate Chain
A governance pipeline that validates proposals before they become active rules. Accessed via gate_check, apply_proposal, and activate_wrapper.
Safety Gate: Blocks proposals that could harm data integrity or violate compliance constraints
Cost Gate: Estimates the cost of the new rule (e.g., additional API calls) and warns if high
Regression Gate: Runs A/B experiments to ensure the new rule doesn't break passing cases
Usefulness Gate: Verifies the rule actually solves the reported problem
When all gates ALLOW, the proposal is automatically activated and appended to the learned rules.
High-risk or gate-blocked proposals require human approval (pending_approval=true).