Design Review Package · GoF / SOLID / TDD

Agent Self-Improvement System

A flexible, interface-first architecture for improving the agent wrapper around cheap LLMs: system messages, prompts, tools, workflows, context, evaluations, memory, and rollback-safe improvement proposals.

Program to Interfaces Abstract Factories Small Testable Objects TDD First No Pioneer Dependency

Review Map

1. Big Design Change
Split into factory-created families.
2. Abstract Factory Families
Swappable object families.
3. Finer Implementation Phases
Small reviewable steps.
4. Interface Catalog
Responsibilities only, not implementations.
5. First Failing Tests
Phase 00 and Phase 01 tests.
6. Improvement Flow
How a proposed improvement moves safely.

1. Big Design Change: Make the System More Flexible

Recommendation: Break the design into smaller interface groups and create compatible object families with Abstract Factories. The core agent should not know whether it is using SQLite or Postgres, Gemini Flash Lite or another cheap LLM, local files or a memory database, pytest or another evaluator, or one workflow engine versus another.

Before

A single design phase tried to cover storage, prompts, LLM calls, tools, memory, workflow, evaluation, and reports together.

Risk: too much coupling and too many reasons to change.

After

Each subsystem has interfaces, contracts, tests, and a factory. The core system receives a family of compatible objects.

Benefit: we can swap whole subsystems without changing the agent kernel.

2. Abstract Factory Families

The Abstract Factory pattern gives us a way to create related objects without binding the core system to concrete classes.

AgentKernel
  └── uses AgentRuntimeFamily
        ├── PersistenceFactory
        ├── LlmFactory
        ├── PromptFactory
        ├── ToolFactory
        ├── ContextFactory
        ├── WorkflowFactory
        ├── EvaluationFactory
        ├── ImprovementFactory
        └── ReportingFactory

PersistenceFactory

Creates repositories, unit-of-work objects, schema managers, and migration runners.

Abstract FactoryRepositoryUnit of Work

LlmFactory

Creates cheap LLM clients, retry decorators, token counters, budget guards, and response parsers.

Abstract FactoryAdapterDecorator

PromptFactory

Creates system-message builders, prompt composers, templates, validators, and version diff tools.

Abstract FactoryBuilderStrategy

ToolFactory

Creates tool registries, tool adapters, permission policies, command wrappers, and audit loggers.

Abstract FactoryCommandAdapter

ContextFactory

Creates context selectors, memory readers, memory writers, note managers, and summarizers.

Abstract FactoryStrategyMemento

EvaluationFactory

Creates evaluators, verdict policies, scoring rules, regression checks, and safety gates.

Abstract FactoryChain of ResponsibilityStrategy

3. Finer Implementation Phases

These phases are intentionally small so the team can review interfaces before implementation. Each phase starts with failing tests.

PhaseGoalInterfaces / ClassesTests to Write FirstWhat Fails FirstDone Means
00
Vocabulary + IDs
Define common terms, IDs, and immutable value objects. AgentId, RunId, VersionId, PromptId, ToolId, EvaluationId, ImprovementId IDs are stable, printable, comparable, and cannot be empty. Value object constructors not implemented. Every subsystem can share IDs without using raw strings everywhere.
01
Database Contracts
Design the storage schema and repository contracts. ISchemaManager, IRunRepository, IPromptVersionRepository, IToolVersionRepository, IEvaluationRepository, IImprovementRepository Schema creates required tables; repositories save and retrieve versions; unit of work commits/rolls back. NotImplementedError from schema manager and repositories. Tests describe the storage contract independent of SQLite/Postgres.
02
Factory Contracts
Create Abstract Factory interfaces for all object families. IAgentRuntimeFamilyFactory, IPersistenceFactory, ILlmFactory, IPromptFactory, IToolFactory, IContextFactory, IWorkflowFactory, IEvaluationFactory Factory creates a complete family; family validates compatibility; missing factories fail fast. Factory methods raise NotImplementedError. AgentKernel depends only on factory interfaces.
03
Composition Root
Centralize object construction and dependency injection. CompositionRoot, RuntimeProfile, RuntimeFamilyValidator Given a runtime profile, CompositionRoot builds AgentKernel with all required ports. build_kernel() raises NotImplementedError. Only CompositionRoot knows concrete classes.
04
LLM Engine Port
Hide cheap LLM vendors behind an interface. ILlmClient, ILlmRequest, ILlmResponse, ITokenCounter, IBudgetGuard, IResponseParser LLM request includes model, messages, tools, budget, trace ID; response captures tokens and raw text. Concrete Gemini/OpenAI/local adapters not implemented. Core can swap LLM providers with no workflow changes.
05
Prompt + System Message Versioning
Make prompts and system messages versioned, testable artifacts. ISystemMessageBuilder, IPromptComposer, IPromptTemplateRepository, IPromptValidator, IPromptDiff Prompt version includes parent version, rationale, author, diff, and active/inactive state. Prompt builder and repository methods not implemented. Every prompt/system-message change is traceable and reversible.
06
Tool Contract Layer
Make tools swappable and permissioned. ITool, IToolRegistry, IToolAdapter, IToolPermissionPolicy, IToolExecutionLogger Tool calls are validated before execution; denied tools never run; all calls are logged. Tool registry and permission policy methods not implemented. Agent can use different tool sets safely.
07
Context Selection
Control which memory, notes, files, and examples enter the prompt. IContextSelector, IContextBudgeter, IMemoryReader, INoteReader, IContextPackBuilder Selector respects token budget and ranking policy; context pack is reproducible. Context selectors raise NotImplementedError. Context usage can be improved without touching LLM code.
08
Workflow Skeletons
Define workflows as testable templates. IWorkflow, WorkflowTemplate, Step, StepResult, IWorkflowRunner Runner executes steps in order; failed step stops or branches according to policy. WorkflowTemplate.run() not implemented. Different workflows can be swapped by Strategy/Template Method.
09
Evaluation Layer
Measure output quality and regressions. IEvaluator, ITestSuiteRunner, IGoldenCaseRepository, IScoreCard, IRegressionDetector Evaluator produces structured scorecards; regression detector compares against baseline. Evaluator returns NotImplementedError. Every change can be judged against evidence.
10
Verdict + Gates
Approve or reject changes using explicit rules. IVerdictPolicy, IGate, GateChain, SafetyGate, CostGate, RegressionGate Gate chain stops on hard failure; verdict explains reasons. Gate chain methods not implemented. No improvement can activate without passing gates.
11
Improvement Proposals
Represent proposed changes as first-class objects. IImprovementProposal, IProposalGenerator, IProposalRepository, IChangeSet, IRationale Proposal includes target artifact, change type, expected benefit, risk, tests required. Proposal generator not implemented. Proposed wrapper improvements are reviewable before activation.
12
Experiment Runner
Run baseline vs candidate comparisons. IExperimentRunner, ABRunConfig, BaselineRun, CandidateRun, ExperimentResult Same inputs run against baseline and candidate; results are comparable. Experiment runner not implemented. We can prove a wrapper change improved behavior before activation.
13
Activation + Rollback
Safely activate approved changes and roll back bad ones. IActivationService, IRollbackService, ISnapshotStore, ActiveVersionRegistry Activation creates snapshot; rollback restores previous active versions. Snapshot and rollback services not implemented. No change is permanent without a recoverable previous version.
14
Memory + Notes Improvement
Improve agent memory/notes over time without poisoning context. IMemoryWriter, INoteWriter, IMemoryPolicy, INotePolicy, IStalenessDetector Memory changes require source, reason, and expiration/staleness policy. Policies not implemented. Memory improves usefulness without becoming junk drawer context.
15
Plugin Registry
Allow new implementations without editing the kernel. IPluginRegistry, IRuntimeProfileRepository, IFactoryResolver Registry resolves factories by profile; unknown profile fails with useful error. Factory resolver not implemented. New vendors/backends/workflows can be added as plugins.
16
Reports + Review UI
Present decisions, experiments, and active versions clearly. IReportBuilder, IDesignReviewPageBuilder, IRunReportRepository Report includes baseline, candidate, verdict, approved changes, rollback info. Report builders not implemented. Team can review system behavior without reading raw logs.

4. Proposed Interfaces and Responsibilities

Kernel Interfaces

  • IAgentKernel — runs one task using injected subsystems.
  • IAgentRuntimeFamily — bundle of compatible factories.
  • ICompositionRoot — builds the concrete runtime from a profile.
  • IRuntimeProfile — describes selected DB, LLM, workflow, tool, context, and evaluation family.

Persistence Interfaces

  • ISchemaManager — creates and migrates required tables.
  • IUnitOfWork — commits or rolls back repository changes.
  • IRunRepository — stores agent run records and trace IDs.
  • IVersionRepository — stores prompt, system-message, workflow, and tool versions.

LLM Interfaces

  • ILlmClient — sends a vendor-neutral LLM request.
  • ILlmRequestBuilder — builds messages, tools, model, and budget.
  • ITokenCounter — estimates or records token usage.
  • IBudgetGuard — blocks requests that exceed configured cost limits.

Prompt Interfaces

  • ISystemMessageBuilder — creates versioned system messages.
  • IPromptComposer — combines task, context, tools, examples, and rules.
  • IPromptValidator — checks required sections and forbidden content.
  • IPromptDiff — explains differences between prompt versions.

Tool Interfaces

  • ITool — one executable tool contract.
  • IToolRegistry — lists and resolves available tools.
  • IToolPermissionPolicy — approves or denies tool execution.
  • IToolExecutionLogger — records request, result, and error.

Context + Memory Interfaces

  • IContextSelector — chooses relevant memory, notes, examples, and files.
  • IContextBudgeter — enforces token budget.
  • IMemoryReader / IMemoryWriter — reads and writes long-lived memory.
  • IStalenessDetector — finds old or risky context.

Evaluation Interfaces

  • IEvaluator — scores a run or output.
  • ITestSuiteRunner — runs unit/golden/regression tests.
  • IVerdictPolicy — converts evidence into approve/reject/manual-review.
  • IGate — one rule in the approval chain.

Improvement Interfaces

  • IProposalGenerator — proposes wrapper improvements.
  • IChangeSet — describes exact changes to prompts/tools/workflows/context.
  • IExperimentRunner — compares baseline against candidate.
  • IActivationService / IRollbackService — activates or restores versions.

5. GoF Pattern Map

PatternWhere It FitsWhy We Use It
Abstract FactoryRuntime families: persistence, LLM, prompts, tools, context, workflow, evaluation.Swap compatible object families without changing AgentKernel.
StrategyContext selection, LLM choice, evaluation rules, improvement mutation methods.Change algorithms independently.
AdapterGemini/OpenAI/local LLM clients, Letta-style agents, tool APIs.Convert external APIs into our internal ports.
DecoratorLLM logging, retry, cost tracking, cache, rate-limit wrappers.Add behavior without subclass explosions.
CommandTool calls, proposed changes, activation, rollback.Represent actions as objects that can be logged, replayed, approved, or rejected.
Chain of ResponsibilitySafety gate, cost gate, regression gate, human-review gate.Stop unsafe changes early and explain why.
Template MethodWorkflow skeletons.Keep the workflow shape stable while allowing steps to vary.
MementoSnapshots before activation.Rollback to previous active versions.
ObserverRun events, evaluation events, report updates.Allow logging/reporting without coupling to execution.

6. First Failing Unit Tests

The first tests should fail by design. They document the contract before implementation.

Important: These are not tests that assert NotImplementedError. They call the future behavior directly, so pytest fails until the implementation exists.

Phase 00 / 02 Factory Contract Test Example

def test_runtime_family_factory_creates_complete_family():
    factory = NotImplementedRuntimeFamilyFactory()

    family = factory.create_runtime_family("sqlite_gemini_flash_lite_default")

    assert family.persistence_factory is not None
    assert family.llm_factory is not None
    assert family.prompt_factory is not None
    assert family.tool_factory is not None
    assert family.context_factory is not None
    assert family.workflow_factory is not None
    assert family.evaluation_factory is not None
    assert family.improvement_factory is not None

Phase 01 Database Contract Test Example

def test_schema_manager_creates_core_tables():
    schema_manager = NotImplementedSchemaManager()

    schema_manager.create_schema()

    assert schema_manager.has_table("agent_runs")
    assert schema_manager.has_table("prompt_versions")
    assert schema_manager.has_table("system_message_versions")
    assert schema_manager.has_table("tool_versions")
    assert schema_manager.has_table("workflow_versions")
    assert schema_manager.has_table("evaluations")
    assert schema_manager.has_table("improvement_proposals")
    assert schema_manager.has_table("activation_snapshots")

7. Safe Self-Improvement Flow

The system does not let the LLM rewrite itself directly. It proposes changes, tests them, gates them, activates them, and keeps rollback snapshots.

User / Schedule
  → AgentKernel runs task
  → RunRepository stores trace
  → Evaluator scores result
  → ProposalGenerator suggests wrapper improvement
  → ExperimentRunner compares baseline vs candidate
  → GateChain checks safety, cost, regression, usefulness
  → VerdictPolicy approves / rejects / sends to manual review
  → ActivationService activates approved version
  → SnapshotStore keeps rollback point
  → ReportingService creates team review report

8. Suggested Package Layout

agent_self_improvement/
  core/
    kernel.py
    composition_root.py
    runtime_profile.py
  contracts/
    ids.py
    factories.py
    persistence.py
    llm.py
    prompts.py
    tools.py
    context.py
    workflow.py
    evaluation.py
    improvement.py
    reporting.py
  implementations/
    sqlite_persistence/
    gemini_flash_lite_llm/
    default_prompting/
    local_tools/
    pytest_evaluation/
  tests/
    test_phase_00_value_objects.py
    test_phase_00_factory_contracts.py
    test_phase_01_database_contracts.py
    test_phase_02_composition_root.py

9. Design Review Decision Points

Approve Now

  • Use Abstract Factory families.
  • Keep AgentKernel free of concrete implementations.
  • Start with failing tests for value objects, factories, and database contracts.
  • Record every run, change, evaluation, verdict, activation, and rollback.

Do Not Implement Yet

  • Do not wire real LLM calls yet.
  • Do not build recursive automatic activation yet.
  • Do not let proposed changes become active without gates.
  • Do not hard-code SQLite, Gemini, or any one tool system into the core.