/home/adamsl/Archon),
and an engineering assessment of what our Letta-based self-improving agent system (Mazda and friends)
can reuse, adapt, or integrate from it.
Archon is a self-hostable, governed agentic automation engine. It runs multi-step workflows that mix deterministic steps (bash / TypeScript / Python scripts) with AI agent steps (Claude Code SDK, Codex SDK, and others), with human approval gates and a full audit trail — driven from Slack, Telegram, GitHub, Discord, the web UI, or the CLI. Its most mature use is agentic coding against git repos; the same engine is being extended to general business-operations automation.
A workflow executor that reads YAML DAG definitions, runs nodes in topological order (independent nodes concurrently), substitutes outputs between nodes, pauses at approval gates, and records every step transition to a database.
AI backends behind one interface (IAgentProvider): Claude (Agent SDK),
Codex, Pi (~20 LLM backends), OpenCode, and
Copilot. Providers are registered in a typed registry and selected per workflow or per node.
Platform adapters behind one interface (IPlatformAdapter): Web UI (SSE streaming),
Slack, Telegram, GitHub (webhooks + @archon mentions), Discord, and a CLI that works
without the server at all.
Stack: Bun workspaces monorepo — @archon/paths, @archon/git,
@archon/providers, @archon/isolation, @archon/workflows (the engine),
@archon/core, @archon/adapters, @archon/server (Hono + OpenAPI),
@archon/web (React), @archon/cli. Storage is SQLite by default
(~/.archon/archon.db, zero setup) or PostgreSQL via DATABASE_URL.
License is MIT.
The CLI runs workflows directly, no server needed. It must be run from inside a git repository
(subdirectories work). From /home/adamsl/Archon:
# Dev server + web UI (server :3090, web :5173)
bun run dev
# Everything below works standalone from any git repo:
bun run cli workflow list # what workflows exist here (bundled + repo + ~/.archon)
bun run cli workflow run assist "What does the orchestrator do?"
bun run cli workflow run implement "Add auth" # auto-creates an isolated worktree + branch
bun run cli workflow run implement --branch feature-auth "Add auth"
bun run cli workflow run quick-fix --no-worktree "Fix typo" # opt out of isolation
bun run cli workflow run implement "Add auth" --detach # background child process
bun run cli workflow status # active runs (running + paused)
bun run cli workflow runs --json # recent runs, machine-readable
bun run cli workflow get <run-id> --verbose # one run, per-node summary
bun run cli workflow resume <run-id> # re-run a failed run, skipping completed nodes
bun run cli workflow abandon <run-id> # discard a non-terminal run
# Approval gates (a paused run waits for one of these):
bun run cli workflow approve <run-id> "looks good, also rename X"
bun run cli workflow reject <run-id> "wrong file, target server.py"
bun run cli validate workflows # lint all workflow YAML + referenced resources
bun run cli isolation list # active worktrees
bun run cli isolation cleanup --merged # remove worktrees whose branches merged
bun run cli doctor # verify setup (claude binary, gh auth, DB, adapters)
In chat surfaces (Slack/Telegram/web) the same lifecycle is exposed as slash commands:
/workflow list, /workflow run <name> <args>, /workflow status,
/workflow approve <id>, /workflow resume <id>, etc. Free-text messages are
routed by an AI router to the best-matching workflow (falling back to archon-assist).
| Scope | Path | Notes |
|---|---|---|
| Bundled defaults | .archon/workflows/defaults/ in the Archon repo | ~20 ready-made workflows: archon-feature-development, archon-fix-github-issue, archon-plan-to-pr, archon-piv-loop, archon-ralph-dag… |
| Per-repo | <repo>/.archon/workflows/*.yaml | Overrides bundled by name; discovered recursively at runtime |
| Home / global | ~/.archon/workflows/ | Applies to every project; priority bundled < global < project |
A workflow is a YAML file with a nodes: list. Each node has an id, exactly one
node-type key, and optional depends_on edges. Nodes in the same topological layer run
concurrently. Everything is Zod-validated at load time (cycles, unknown deps, bad
$nodeId.output references, unknown providers all rejected before anything runs).
| Type | What it does | AI? |
|---|---|---|
prompt: | Inline prompt sent to the configured AI provider | Yes |
command: | Runs a named command file from .archon/commands/ (a reusable prompt template) | Yes |
loop: | Iterative AI prompt repeated until a completion signal; supports fresh_context and $LOOP_PREV_OUTPUT | Yes |
bash: | Shell script; stdout captured as $nodeId.output; receives managed per-project env vars | No |
script: | Inline or named TypeScript (runtime: bun) or Python (runtime: uv) with deps: and timeout:; stdout captured | No |
approval: | Human gate — run pauses until approve/reject; capture_response: true stores the human's comment as the node output; rejection feeds $REJECTION_REASON into an on_reject prompt | No |
$nodeId.output — cleaned output of a completed upstream node.$nodeId.output.field — field access when the producer declared output_format
(a JSON schema). Access is strict: an undeclared field fails the consuming node instead of
silently passing ''.output_format — structured JSON output, SDK-grammar-enforced on Claude/Codex/OpenCode and
best-effort (validate + re-ask up to 3×) on Pi/Copilot. A node that declares a schema but can't produce
valid output fails loudly.when: conditions and trigger_rule join semantics for conditional branches.$1..$n, $ARGUMENTS, $ARTIFACTS_DIR (per-run artifacts
directory, pre-created), $WORKFLOW_ID, $BASE_BRANCH, $LOOP_USER_INPUT,
$REJECTION_REASON.name: archon-feature-development
description: |
Use when: Implementing a feature from an existing plan.
nodes:
- id: implement
command: archon-implement # named prompt template
provider: claude
model: large # resolves through the tier system
- id: create-pr
command: archon-create-pr
depends_on: [implement]
context: fresh # new session, not a continuation
- id: verify-pr-base
bash: | # deterministic step, no AI
set -euo pipefail
HEAD_BRANCH=$(git rev-parse --abbrev-ref HEAD)
PR_NUMBER=$(gh pr list --head "$HEAD_BRANCH" --state open --json number -q '.[0].number')
...
depends_on: [create-pr]
name: finance-doc-pipeline
nodes:
- id: parse
prompt: "Parse the statement at $1 into transactions."
output_format: # schema-validated JSON out
type: object
properties:
vendor: { type: string }
total: { type: number }
- id: duplicate-check
script:
runtime: uv # real Python, real DB, deterministic
name: duplicate_lookup # .archon/scripts/duplicate_lookup.py
depends_on: [parse]
- id: human-gate
approval: "Store expense for $parse.output.vendor, total $parse.output.total?"
capture_response: true
depends_on: [duplicate-check]
- id: store
script: { runtime: uv, name: store_expense }
depends_on: [human-gate]
Every workflow run (and optionally every conversation) executes in its own git worktree under
~/.archon/workspaces/<owner>/<repo>/worktrees/, with auto-generated branch names and
deterministic per-worktree dev ports. Parallel runs can't stomp each other. Cleanup commands understand
"branch merged" and "PR closed". Non-destructive by default; it refuses to remove worktrees with
uncommitted changes.
approval: nodes pause the run in the database. Approve/reject arrives later from
any surface (CLI, Slack, web) and resumes execution. Rejections carry a reason into
on_reject prompts; approvals can carry user text into the next loop iteration
($LOOP_USER_INPUT). This is a durable, replayable human-in-the-loop primitive.
Two tables — workflow_runs (run state) and workflow_events (step-level log:
transitions, artifacts, errors) — plus JSONL file logs and a per-run $ARTIFACTS_DIR outside
the repo. Nodes can declare output_type to get typed output sidecars
(nodes/<id>.md + .meta.json) so later runs find outputs by type, not filename
guessing.
AI sessions are immutable; transitions create new linked sessions with an explicit reason
(audit trail via parent_session_id). persist_session: true on a node keeps one
provider session across runs, keyed by (workflow, node, scope) — long-lived per-node memory.
Workflows say model: small|medium|large or @custom-alias; config maps tiers
to concrete provider/model/effort (globally, per repo, or per user). Swapping the entire fleet from
Claude to Codex to a cheap OpenRouter model is a config edit, not a workflow edit. This is
"wrapper-not-model" made operational.
Per-user AI credentials and GitHub tokens encrypted at rest (AES-256-GCM); runs execute with the
acting user's credentials. Per-node allowed_tools/denied_tools
restrictions, per-node MCP server configs, hooks, and sub-agent definitions (Claude). Workflows can
declare requires: [github] to hard-block before any cost is incurred.
Two YAML files, merged (repo overrides global):
| File | Purpose |
|---|---|
~/.archon/config.yaml | Install-wide: default assistant, per-provider model defaults, tiers, aliases |
<repo>/.archon/config.yaml | Repo-specific overrides + env:, docs path, defaults opt-out |
assistants:
claude:
model: sonnet # or opus / haiku / claude-* / inherit
codex:
model: gpt-5.3-codex
modelReasoningEffort: medium
tiers: # what small/medium/large mean HERE
small: { provider: claude, model: haiku }
medium: { provider: claude, model: sonnet }
large: { provider: codex, model: gpt-5.5, effort: high }
aliases:
"@cheap": { provider: pi, model: openrouter/qwen/qwen3-coder }
Priority: workflow-level options → config defaults → SDK defaults. Per-project env vars
(codebase_env_vars table, managed in the web UI) are injected into Claude/Codex/bash/script
node subprocess environments.
Our system (per Mazda — A Developer's Manual and the Self-Improvement Plan) has two faces:
the runtime — Mazda, a Letta orchestrator agent, delegating to five minion agents that
drive Claude Agent SDK sessions against a finance MCP server — and the control plane —
the interface-first Python agent_self_improvement package (trace → judge → propose → A/B →
gate → activate → rollback; nine factory families; "improve the wrapper, never the model").
| Dimension | Archon | Our Letta/Mazda system |
|---|---|---|
| Core unit of work | A workflow run: a YAML DAG executed by a stateless engine; state lives in DB rows + artifacts | A stateful agent: Letta server holds memory blocks, tools, and conversation history per agent |
| Orchestration | Deterministic engine orders the steps; AI only fills in steps (plus an AI router for chat entry) | An LLM (Mazda) decides ordering/delegation at runtime; workflow order is prompt-encoded |
| Memory / learning | None built in — sessions can persist per node, but there is no self-editing memory and no improvement loop | The whole point: memory self-editing is live and proven; the gated improvement loop is built (not yet wired) |
| Human gates | First-class approval: nodes, durable across restarts, actionable from any surface |
Designed (GateChain in the self-improvement plan) but not implemented as a runtime primitive |
| Audit / evidence | workflow_runs + workflow_events + JSONL logs + per-run artifact dirs, out of the box |
Trace capture is a planned service (TraceCommandService); today evidence is ad-hoc (API reads, dashboards) |
| Deterministic steps | bash: and script: nodes (bun/uv Python) as peers of AI nodes in the same DAG |
Deterministic services live behind the finance MCP server / Python services; the agent must choose to call them |
| AI backends | 5 providers behind IAgentProvider (Claude, Codex, Pi≈20 backends, OpenCode, Copilot); per-node override |
Letta-managed models + minions hard-wired to the Claude Agent SDK executor |
| Entry surfaces | Slack, Telegram, GitHub, Discord, Web (SSE), CLI — one adapter interface | Custom dashboard (Python stdlib server + vanilla JS), Telegram for Scissari, letta-code CLI |
| Language / runtime | Bun + strict TypeScript + Zod monorepo | letta-code is also Bun + TypeScript; control plane is Python; dashboard is stdlib Python |
| Storage | Own SQLite (default) or Postgres, 18 remote_agent_* tables |
Letta's Postgres (agent state) + logger API + per-project files |
| Failure posture | Fail fast, resume from last completed node, never auto-kill ambiguous work | Recovery policies exist in letta-code (turn recovery, approval recovery) but per-incident, not engine-level |
| Component | Why it drops in cleanly |
|---|---|
Archon CLI as a dev tool on letta-code and rol_finances |
Zero integration required: bun run cli workflow run implement "…" from either repo gives
us isolated-worktree coding agents, plan→PR pipelines, and PR review workflows today. Both repos are git
repos, which is Archon's only hard requirement. |
@archon/git + @archon/isolation packages |
Leaf packages (no @archon/core dependency), typed worktree/branch/repo operations with
error classification. letta-code is also Bun+TS — these could be imported for Scissari's executor or any
agent that needs safe parallel checkouts. |
Bundled workflows (archon-fix-github-issue, archon-plan-to-pr,
archon-comprehensive-pr-review…) |
Battle-tested YAML for the coding tasks we currently hand-drive; also the best reference corpus for authoring our own pipelines. |
| Approval-gate lifecycle (run/pause/approve/reject/resume, from any surface) | Usable immediately by wrapping finance steps in a Archon workflow — no code changes, just YAML. |
parent_session_id + transition_reason. Applied to Mazda's memory edits, this
would give the self-improvement loop the provenance chain it needs for rollback.output_format +
schema-validated $node.output.field access, failing loudly on misses. This is our
"stable handoff contracts" between stage agents, enforced by the engine instead of by convention.small/medium/large, config
binds them. For the wrapper-improvement thesis this makes "model" just another wrapper knob that A/B
experiments can flip without touching prompts.IWorkflowStore, WorkflowDeps injection).
The engine package has zero DB/AI/config dependencies; everything is injected. Our nine-factory-families
design shares this instinct — Archon proves it works at production scale in TS.| Source | Adaptation | Effort |
|---|---|---|
packages/workflows/ (engine: loader, dag-executor, schemas) |
Use as the execution substrate for the Mazda document pipeline — each stage agent becomes a node.
Consumable as a library thanks to WorkflowDeps injection. |
Low–Medium |
packages/providers/src/community/ (Pi/OpenCode as templates) |
Write a LettaProvider implementing IAgentProvider: map
sendQuery() → POST /v1/agents/<id>/messages on our Letta server
(100.80.49.10:8283), stream chunks back, map session-resume to the agent's persistent conversation.
Community providers (builtIn: false) are the sanctioned extension point — Pi and OpenCode
show the exact shape. |
Medium (~1 package dir + registry entry) |
script: node with runtime: uv |
Run our existing Python deterministic services (duplicate lookup, vendor map, expense repository) as first-class DAG nodes — no rewrite, they already speak Python. | Low |
workflow_events schema + event emitter |
Adopt as the TraceCommandService implementation for the improvement loop: every stage transition, artifact, and error is already an event row. | Low |
Approval-gate + $REJECTION_REASON machinery |
Implements the improvement loop's gate step: "propose → human gate → activate" becomes a
three-node Archon workflow with rollback as an on_reject script node. |
Low |
Out of the box Archon can't talk to a Letta agent — AI nodes drive stateless SDK sessions
(Claude/Codex/…). Using Mazda's minions from Archon requires the LettaProvider above,
or keeping minions as Claude-SDK sessions and letting Letta-Mazda sit outside the engine.
Letta agents remember everything server-side; Archon assumes fresh sessions per node
(persist_session is per-(workflow, node, scope), not a global agent memory). Mapping
Mazda's memory-block learning onto Archon nodes needs a deliberate decision about which state lives
where — don't let both sides think they own history.
The CLI and worktree isolation require a git repo cwd. Fine for rol_finances and
letta-code, but finance document runs don't need worktrees — use
--no-worktree or the server API to avoid pointless checkout churn.
Archon writes its own SQLite/Postgres (18 remote_agent_* tables); Letta has its own
Postgres. That's fine (single-tenant, two concerns) but means run-evidence joins across the two need an
export step — e.g. a bash: node posting run summaries to our logger API (:8284).
Archon will not judge, propose, or A/B anything. The agent_self_improvement control
plane remains ours; Archon is only the trustworthy executor + gatekeeper underneath it.
Archon is a fast-moving codebase (Bun-specific test isolation quirks, strict Zod/TS conventions, generated bundles that must be regenerated). Pin a version for integration; treat upgrades as scheduled maintenance. MIT license poses no constraint.
| Option | Verdict | Rationale |
|---|---|---|
| 1 · Use parts of Archon as-is | YES — start now | Zero-risk wins: run Archon CLI workflows (fix-issue, plan-to-PR, PR review) against
letta-code and rol_finances for our own development work. This also builds
operational familiarity before any deeper integration. |
| 2 · Adapt Archon's concepts / code | YES — the main prize | Adopt the engine (YAML DAG + approval gates + workflow_events) as the execution substrate the Mazda
plans keep specifying but haven't built: stage order as data, deterministic Python services as
script: nodes, human gates and audit for free. |
| 3 · Integrate into the Letta system | SELECTIVE | Integrate at two seams only: (a) a community LettaProvider so DAG nodes can be Letta
agents, and (b) an evidence bridge from workflow_events to our logger/improvement loop.
Do not try to merge databases, replace the dashboard, or move Mazda's memory into Archon —
the learning loop stays ours. |
rol_finances:
bun run cli workflow list, then run archon-assist and one coding workflow in a
worktree. Confirm ~/.archon/ workspace layout, artifacts, and the approve/reject loop from
the CLI.rol_finances/.archon/workflows/finance-doc-pipeline.yaml,
with the deterministic steps as script: runtime: uv nodes calling our existing Python
services, and output_format schemas as the stage contracts. Run with
--no-worktree.LettaProvider in packages/providers/src/community/letta/
(modeled on the Pi provider, registered builtIn: false): sendQuery() →
Letta REST messages API, streaming chunks back; session-resume → the agent's own conversation. Then a
DAG node can be Mazda or a minion via provider: letta.bash:/script: node (or a
small consumer of workflow_events) that posts run summaries to the logger API
(:8284) so improvement-loop judging can score Archon runs alongside live-agent traces.agent_self_improvement Python services. This is the first live use of the gated loop that
§4 of the Self-Improvement Plan says is built-but-not-wired.GET /api/workflows/runs, server on :3090) the same way we poll Letta — runs, statuses, and
pending approvals surfaced next to agent health./home/adamsl/Archon (branch dev) — architecture reference in its CLAUDE.md; docs site content under packages/docs-web/./home/adamsl/Archon/.archon/workflows/defaults/.