Review Map
1. Big Design Change: Make the System More Flexible
Before
A single design phase tried to cover storage, prompts, LLM calls, tools, memory, workflow, evaluation, and reports together.
Risk: too much coupling and too many reasons to change.
After
Each subsystem has interfaces, contracts, tests, and a factory. The core system receives a family of compatible objects.
Benefit: we can swap whole subsystems without changing the agent kernel.
2. Abstract Factory Families
The Abstract Factory pattern gives us a way to create related objects without binding the core system to concrete classes.
AgentKernel
└── uses AgentRuntimeFamily
├── PersistenceFactory
├── LlmFactory
├── PromptFactory
├── ToolFactory
├── ContextFactory
├── WorkflowFactory
├── EvaluationFactory
├── ImprovementFactory
└── ReportingFactoryPersistenceFactory
Creates repositories, unit-of-work objects, schema managers, and migration runners.
Abstract FactoryRepositoryUnit of WorkLlmFactory
Creates cheap LLM clients, retry decorators, token counters, budget guards, and response parsers.
Abstract FactoryAdapterDecoratorPromptFactory
Creates system-message builders, prompt composers, templates, validators, and version diff tools.
Abstract FactoryBuilderStrategyToolFactory
Creates tool registries, tool adapters, permission policies, command wrappers, and audit loggers.
Abstract FactoryCommandAdapterContextFactory
Creates context selectors, memory readers, memory writers, note managers, and summarizers.
Abstract FactoryStrategyMementoEvaluationFactory
Creates evaluators, verdict policies, scoring rules, regression checks, and safety gates.
Abstract FactoryChain of ResponsibilityStrategy3. Finer Implementation Phases
These phases are intentionally small so the team can review interfaces before implementation. Each phase starts with failing tests.
| Phase | Goal | Interfaces / Classes | Tests to Write First | What Fails First | Done Means |
|---|---|---|---|---|---|
| 00 Vocabulary + IDs |
Define common terms, IDs, and immutable value objects. | AgentId, RunId, VersionId, PromptId, ToolId, EvaluationId, ImprovementId | IDs are stable, printable, comparable, and cannot be empty. | Value object constructors not implemented. | Every subsystem can share IDs without using raw strings everywhere. |
| 01 Database Contracts |
Design the storage schema and repository contracts. | ISchemaManager, IRunRepository, IPromptVersionRepository, IToolVersionRepository, IEvaluationRepository, IImprovementRepository | Schema creates required tables; repositories save and retrieve versions; unit of work commits/rolls back. | NotImplementedError from schema manager and repositories. | Tests describe the storage contract independent of SQLite/Postgres. |
| 02 Factory Contracts |
Create Abstract Factory interfaces for all object families. | IAgentRuntimeFamilyFactory, IPersistenceFactory, ILlmFactory, IPromptFactory, IToolFactory, IContextFactory, IWorkflowFactory, IEvaluationFactory | Factory creates a complete family; family validates compatibility; missing factories fail fast. | Factory methods raise NotImplementedError. | AgentKernel depends only on factory interfaces. |
| 03 Composition Root |
Centralize object construction and dependency injection. | CompositionRoot, RuntimeProfile, RuntimeFamilyValidator | Given a runtime profile, CompositionRoot builds AgentKernel with all required ports. | build_kernel() raises NotImplementedError. | Only CompositionRoot knows concrete classes. |
| 04 LLM Engine Port |
Hide cheap LLM vendors behind an interface. | ILlmClient, ILlmRequest, ILlmResponse, ITokenCounter, IBudgetGuard, IResponseParser | LLM request includes model, messages, tools, budget, trace ID; response captures tokens and raw text. | Concrete Gemini/OpenAI/local adapters not implemented. | Core can swap LLM providers with no workflow changes. |
| 05 Prompt + System Message Versioning |
Make prompts and system messages versioned, testable artifacts. | ISystemMessageBuilder, IPromptComposer, IPromptTemplateRepository, IPromptValidator, IPromptDiff | Prompt version includes parent version, rationale, author, diff, and active/inactive state. | Prompt builder and repository methods not implemented. | Every prompt/system-message change is traceable and reversible. |
| 06 Tool Contract Layer |
Make tools swappable and permissioned. | ITool, IToolRegistry, IToolAdapter, IToolPermissionPolicy, IToolExecutionLogger | Tool calls are validated before execution; denied tools never run; all calls are logged. | Tool registry and permission policy methods not implemented. | Agent can use different tool sets safely. |
| 07 Context Selection |
Control which memory, notes, files, and examples enter the prompt. | IContextSelector, IContextBudgeter, IMemoryReader, INoteReader, IContextPackBuilder | Selector respects token budget and ranking policy; context pack is reproducible. | Context selectors raise NotImplementedError. | Context usage can be improved without touching LLM code. |
| 08 Workflow Skeletons |
Define workflows as testable templates. | IWorkflow, WorkflowTemplate, Step, StepResult, IWorkflowRunner | Runner executes steps in order; failed step stops or branches according to policy. | WorkflowTemplate.run() not implemented. | Different workflows can be swapped by Strategy/Template Method. |
| 09 Evaluation Layer |
Measure output quality and regressions. | IEvaluator, ITestSuiteRunner, IGoldenCaseRepository, IScoreCard, IRegressionDetector | Evaluator produces structured scorecards; regression detector compares against baseline. | Evaluator returns NotImplementedError. | Every change can be judged against evidence. |
| 10 Verdict + Gates |
Approve or reject changes using explicit rules. | IVerdictPolicy, IGate, GateChain, SafetyGate, CostGate, RegressionGate | Gate chain stops on hard failure; verdict explains reasons. | Gate chain methods not implemented. | No improvement can activate without passing gates. |
| 11 Improvement Proposals |
Represent proposed changes as first-class objects. | IImprovementProposal, IProposalGenerator, IProposalRepository, IChangeSet, IRationale | Proposal includes target artifact, change type, expected benefit, risk, tests required. | Proposal generator not implemented. | Proposed wrapper improvements are reviewable before activation. |
| 12 Experiment Runner |
Run baseline vs candidate comparisons. | IExperimentRunner, ABRunConfig, BaselineRun, CandidateRun, ExperimentResult | Same inputs run against baseline and candidate; results are comparable. | Experiment runner not implemented. | We can prove a wrapper change improved behavior before activation. |
| 13 Activation + Rollback |
Safely activate approved changes and roll back bad ones. | IActivationService, IRollbackService, ISnapshotStore, ActiveVersionRegistry | Activation creates snapshot; rollback restores previous active versions. | Snapshot and rollback services not implemented. | No change is permanent without a recoverable previous version. |
| 14 Memory + Notes Improvement |
Improve agent memory/notes over time without poisoning context. | IMemoryWriter, INoteWriter, IMemoryPolicy, INotePolicy, IStalenessDetector | Memory changes require source, reason, and expiration/staleness policy. | Policies not implemented. | Memory improves usefulness without becoming junk drawer context. |
| 15 Plugin Registry |
Allow new implementations without editing the kernel. | IPluginRegistry, IRuntimeProfileRepository, IFactoryResolver | Registry resolves factories by profile; unknown profile fails with useful error. | Factory resolver not implemented. | New vendors/backends/workflows can be added as plugins. |
| 16 Reports + Review UI |
Present decisions, experiments, and active versions clearly. | IReportBuilder, IDesignReviewPageBuilder, IRunReportRepository | Report includes baseline, candidate, verdict, approved changes, rollback info. | Report builders not implemented. | Team can review system behavior without reading raw logs. |
4. Proposed Interfaces and Responsibilities
Kernel Interfaces
- IAgentKernel — runs one task using injected subsystems.
- IAgentRuntimeFamily — bundle of compatible factories.
- ICompositionRoot — builds the concrete runtime from a profile.
- IRuntimeProfile — describes selected DB, LLM, workflow, tool, context, and evaluation family.
Persistence Interfaces
- ISchemaManager — creates and migrates required tables.
- IUnitOfWork — commits or rolls back repository changes.
- IRunRepository — stores agent run records and trace IDs.
- IVersionRepository — stores prompt, system-message, workflow, and tool versions.
LLM Interfaces
- ILlmClient — sends a vendor-neutral LLM request.
- ILlmRequestBuilder — builds messages, tools, model, and budget.
- ITokenCounter — estimates or records token usage.
- IBudgetGuard — blocks requests that exceed configured cost limits.
Prompt Interfaces
- ISystemMessageBuilder — creates versioned system messages.
- IPromptComposer — combines task, context, tools, examples, and rules.
- IPromptValidator — checks required sections and forbidden content.
- IPromptDiff — explains differences between prompt versions.
Tool Interfaces
- ITool — one executable tool contract.
- IToolRegistry — lists and resolves available tools.
- IToolPermissionPolicy — approves or denies tool execution.
- IToolExecutionLogger — records request, result, and error.
Context + Memory Interfaces
- IContextSelector — chooses relevant memory, notes, examples, and files.
- IContextBudgeter — enforces token budget.
- IMemoryReader / IMemoryWriter — reads and writes long-lived memory.
- IStalenessDetector — finds old or risky context.
Evaluation Interfaces
- IEvaluator — scores a run or output.
- ITestSuiteRunner — runs unit/golden/regression tests.
- IVerdictPolicy — converts evidence into approve/reject/manual-review.
- IGate — one rule in the approval chain.
Improvement Interfaces
- IProposalGenerator — proposes wrapper improvements.
- IChangeSet — describes exact changes to prompts/tools/workflows/context.
- IExperimentRunner — compares baseline against candidate.
- IActivationService / IRollbackService — activates or restores versions.
5. GoF Pattern Map
| Pattern | Where It Fits | Why We Use It |
|---|---|---|
| Abstract Factory | Runtime families: persistence, LLM, prompts, tools, context, workflow, evaluation. | Swap compatible object families without changing AgentKernel. |
| Strategy | Context selection, LLM choice, evaluation rules, improvement mutation methods. | Change algorithms independently. |
| Adapter | Gemini/OpenAI/local LLM clients, Letta-style agents, tool APIs. | Convert external APIs into our internal ports. |
| Decorator | LLM logging, retry, cost tracking, cache, rate-limit wrappers. | Add behavior without subclass explosions. |
| Command | Tool calls, proposed changes, activation, rollback. | Represent actions as objects that can be logged, replayed, approved, or rejected. |
| Chain of Responsibility | Safety gate, cost gate, regression gate, human-review gate. | Stop unsafe changes early and explain why. |
| Template Method | Workflow skeletons. | Keep the workflow shape stable while allowing steps to vary. |
| Memento | Snapshots before activation. | Rollback to previous active versions. |
| Observer | Run events, evaluation events, report updates. | Allow logging/reporting without coupling to execution. |
6. First Failing Unit Tests
The first tests should fail by design. They document the contract before implementation.
NotImplementedError. They call the future behavior directly, so pytest fails until the implementation exists.
Phase 00 / 02 Factory Contract Test Example
def test_runtime_family_factory_creates_complete_family():
factory = NotImplementedRuntimeFamilyFactory()
family = factory.create_runtime_family("sqlite_gemini_flash_lite_default")
assert family.persistence_factory is not None
assert family.llm_factory is not None
assert family.prompt_factory is not None
assert family.tool_factory is not None
assert family.context_factory is not None
assert family.workflow_factory is not None
assert family.evaluation_factory is not None
assert family.improvement_factory is not None
Phase 01 Database Contract Test Example
def test_schema_manager_creates_core_tables():
schema_manager = NotImplementedSchemaManager()
schema_manager.create_schema()
assert schema_manager.has_table("agent_runs")
assert schema_manager.has_table("prompt_versions")
assert schema_manager.has_table("system_message_versions")
assert schema_manager.has_table("tool_versions")
assert schema_manager.has_table("workflow_versions")
assert schema_manager.has_table("evaluations")
assert schema_manager.has_table("improvement_proposals")
assert schema_manager.has_table("activation_snapshots")
7. Safe Self-Improvement Flow
The system does not let the LLM rewrite itself directly. It proposes changes, tests them, gates them, activates them, and keeps rollback snapshots.
User / Schedule
→ AgentKernel runs task
→ RunRepository stores trace
→ Evaluator scores result
→ ProposalGenerator suggests wrapper improvement
→ ExperimentRunner compares baseline vs candidate
→ GateChain checks safety, cost, regression, usefulness
→ VerdictPolicy approves / rejects / sends to manual review
→ ActivationService activates approved version
→ SnapshotStore keeps rollback point
→ ReportingService creates team review report8. Suggested Package Layout
agent_self_improvement/
core/
kernel.py
composition_root.py
runtime_profile.py
contracts/
ids.py
factories.py
persistence.py
llm.py
prompts.py
tools.py
context.py
workflow.py
evaluation.py
improvement.py
reporting.py
implementations/
sqlite_persistence/
gemini_flash_lite_llm/
default_prompting/
local_tools/
pytest_evaluation/
tests/
test_phase_00_value_objects.py
test_phase_00_factory_contracts.py
test_phase_01_database_contracts.py
test_phase_02_composition_root.py
9. Design Review Decision Points
Approve Now
- Use Abstract Factory families.
- Keep AgentKernel free of concrete implementations.
- Start with failing tests for value objects, factories, and database contracts.
- Record every run, change, evaluation, verdict, activation, and rollback.
Do Not Implement Yet
- Do not wire real LLM calls yet.
- Do not build recursive automatic activation yet.
- Do not let proposed changes become active without gates.
- Do not hard-code SQLite, Gemini, or any one tool system into the core.