Skip to content

Repository files navigation

Loop Engineering

Banner

StatusTestsRuntimeDocsLicensePython

Loop engineering is replacing yourself as the person who prompts the agent. You design the system that prompts your agents instead.

Loop Engineering is an early-stage Python runtime for building testable agent loops that plan, act, observe, evaluate, recover, and terminate.

Unlike other frameworks that just provide patterns, Loop Engineering provides a working Python runtime with state machines, budget enforcement, and deterministic gates.

Features | Quickstart | Tutorials | Patterns | CLI | Documentation | Contributing


See It Work (20 Seconds, No API Key)

This is real, unedited output from examples/deterministic_multistep_loop.py - a scripted planner/actor/evaluator that intentionally fails step 2 twice before succeeding, so you can watch the state machine recover without spending a single token:

git clone https://github.com/chillum-codeX/loop-engineering.git
cd loop-engineering && pip install -e .&& python examples/deterministic_multistep_loop.py
 [Planner] Creating plan (call #1)
[Actor] Executing step_1 (call #1)
[Evaluator] Evaluating step_1
[Planner] Revising plan from version 1
[Actor] Executing step_2 (call #2)
[Actor] Step 2: INTENTIONAL FAILURE (attempt 1)
[Planner] Revising plan from version 2
[Actor] Executing step_2 (call #3)
[Actor] Step 2: SUCCESS (attempt 2)
[Evaluator] Evaluating step_2
[Evaluator] Step 2: REJECTING (evaluation 1)
[Planner] Revising plan from version 3
[Actor] Executing step_2 (call #4)
[Actor] Step 2: SUCCESS (attempt 3)
[Evaluator] Evaluating step_2
[Evaluator] Step 2: ACCEPTING (evaluation 2)
[Actor] Executing step_3 (call #5)
[Evaluator] Evaluating step_3
[PASS] Final status is COMPLETED: status=COMPLETED
[PASS] Step 2 failure recorded: step_2_failures=2
[PASS] Recovery executed: recoveries=2
[PASS] All steps VERIFIED_COMPLETED: completed=3/3
[PASS] Execution state progressed: final_state=completed
SUMMARY: 8 passed, 0 failed
✅ ACCEPTANCE TEST PASSED

No mocked provider calls, no hidden setup - a failing step gets caught, retried, re-evaluated, and the loop still reaches COMPLETED with an explicit, inspectable trace of every state transition along the way.


Why Loop Engineering?

Interactive AI tools are excellent for direct collaboration. Recurring workflows add a different problem: somebody must define when the agent runs, what it may spend, how its output is checked, and how interrupted work resumes. Loop engineering makes those controls explicit.

The Problem

  • Repeated manual setup: Stable intent is copied between runs
  • Task-specific scripts: Lifecycle and recovery logic become scattered
  • Ad hoc workflows: Lifecycle state, limits, and recovery can be difficult to inspect or reproduce

The Solution

Loop Engineering provides:

  • State Machine: Explicit states with validated transitions
  • Budget Caps: Hard limits prevent token blowout
  • Deterministic Gates: Rule-based checks BEFORE LLM steps (Stripe Minions pattern)
  • Generator/Evaluator Separation: Different models, temperatures, prompts
  • Human Checkpoints: Preserve engineer control at critical points
  • Recovery Handlers: Automatic retry with escalation
  • Persistence: State survives crashes and restarts

Features

Architecture

Discovery -> Handoff -> Verification
|
v
Scheduling <- Persistence <- Human Checkpoints

Key Components

ComponentDescriptionAnti-Pattern Prevented
State MachineExplicit states, validated transitionsAmnesiac Loop
Budget CapsToken/cost/step limits with trackingRunaway Budget
Deterministic GatesRule-based validation before LLMWishful Thinking
Gen/Eval SeparationDifferent configs for generator/evaluatorEgo Loop
Human CheckpointsMandatory human approvalHuman Absenteeism
RecoveryAutomatic retry with backoffInfinite Retry
PersistenceState survives crashesAmnesiac Loop
Worktree IsolationGit worktrees per taskTangled Loop

Quickstart (5 Minutes)

Installation

git clone https://github.com/chillum-codeX/loop-engineering.git
cd loop-engineering
pip install -e .

The package is not yet published to PyPI. The repository and GitHub release are the supported installation sources for v0.4.1.

Create Your First Loop

# Scaffold a new project
loop-engine init --pattern daily-triage --name my-loop
cd my-loop
# Check readiness
loop-engine audit
# Run the loop (dry run first)
loop-engine run --dry-run
# Execute for real
loop-engine run

Python API

importasynciofromloop_engineimportRuntimeConfig, create_runtimeconfig=RuntimeConfig()
config.discovery.skills_dir=".loop/skills"config.persistence.state_dir=".loop/state"config.persistence.format="sqlite"config.handoff.default_token_budget=100_000config.handoff.default_cost_budget=10.0config.handoff.default_step_budget=50runtime=create_runtime(runtime_config=config, max_iterations=50)
result=asyncio.run(runtime.run())
print(f"Completed: {result.status.name}")
print(f"Tasks completed: {result.tasks_completed}")

The built-in runtime safely validates and persists skill contracts. External actions such as modifying GitHub issues or posting to Slack require an explicit tool adapter; the default runtime does not simulate those side effects.


CLI Commands

loop-engine init - Scaffold Projects

loop-engine init --pattern daily-triage --name my-loop

Creates a complete project structure:

my-loop/
|-- loop.yaml # Configuration
|-- README.md # Documentation
|-- .gitignore # Git ignore rules
`-- .loop/
|-- skills/ # SKILL.md files
|-- state/ # Persistent state
`-- worktrees/ # Git worktrees

loop-engine audit - Readiness Score

$ loop-engine audit --suggest
Audit Results
Score: 108/115
[##################░░] 93%
Categories:
Configuration: 20/20
Structure: 15/15
Skills: 15/15
Documentation: 10/10
Git: 10/10
Safety: 15/15
Checkpoints: 15/15
Activity: 8/15
Suggestions:
Add a .github/workflows/*.yml with a 'schedule:' trigger to prove this loop runs unattended

Scoring is out of 115: the first 100 points are static configuration checks; the last 15 (Activity) are a dynamic check for evidence the loop is actually running - a fresh loop-run-log.md or a scheduled GitHub Actions workflow - not just that the right files exist. See docs/LOOP_READINESS_LEVELS.md for the L0-L3 maturity framework this score maps to, and loop_engine/patterns/registry.yaml for the machine-readable pattern metadata audit and cost both read from.

loop-engine cost - Token Estimator

$ loop-engine cost --pattern pr-babysitter --cadence hourly
Cost Estimate: pr-babysitter
Model: claude-sonnet
Cadence: hourly
Per Run:
Input tokens: 30,000
Output tokens: 20,000
Total tokens: 50,000
Cost: $0.1950
Monthly Estimate:
Runs: 730
Total tokens: 36,500,000
Cost: $142.35

loop-engine validate - Check Configurations

loop-engine validate --strict

Patterns

Loop Engineering documents seven starter patterns. Daily Triage and PR Babysitter currently have runnable starters; the remaining patterns await complete tool adapters and end-to-end validation:

PatternCadenceUse CaseAvg Cost/Run
Daily TriageDailyReview and prioritize tasks$0.15
PR BabysitterPer PRMonitor and review pull requests$0.20
CI SweeperOn failureDiagnose and fix CI failures$0.35
Dependency SweeperWeeklyUpdate and validate dependencies$0.45
Changelog DrafterPer releaseGenerate release notes$0.25
Post-Merge CleanupPost-mergeClean up after merges$0.10
Issue TriageDailyTriage and route issues$0.20

Pattern Structure

Each pattern includes:

  • SKILL.md: Complete specification (WHEN, READ, JUDGE, OUTPUT, STOP)
  • Configuration: Pre-tuned for the pattern
  • Cost Estimates: Heuristic planning estimates; real runs use provider-reported token accounting
  • Safety Guidelines: Budget limits and checkpoints
  • Starter Template: Clone-and-run project

Where It Fits

Loop Engineering is a runtime layer for repeatable agent workflows. It can sit around a model provider or coding agent; it is not a replacement for those tools. Its scope is lifecycle control: explicit state, budgets, deterministic checks, recovery, checkpoints, and persistence.

Provider and product capabilities change quickly, so this project does not claim feature superiority over Claude Code, Codex, Grok, or other agent frameworks. See the tool selection guide for a workflow-oriented comparison.


Verification and Evidence

The repository separates three kinds of evidence:

  • Unit/integration suite: 146 tests covering runtime transitions, budgets, persistence, adapters, CLI behavior, and benchmark evaluators.
  • Deterministic evaluator validation: oracle answers must pass and empty negative controls must fail. Results are stored in experiments/results/deterministic_validation.json.
  • Live provider smoke test: one paid OpenRouter request verified provider-reported token and cost accounting. The response body and API key are not stored; sanitized evidence is in experiments/results/live_provider_smoke_paid.json.

Reproduce the non-secret checks:

python -m pytest tests/ -q
python -m experiments.deterministic_runner
python -m build
python -m twine check dist/*

Historical mock benchmark outputs are not evidence of model quality or security effectiveness. No SOTA, production-readiness, or comparative performance claim is made from them.


Architecture

The Five Phases

  1. Discovery: Load state, discover tasks, build ledger
  2. Handoff: Reserve budget, create worktree, setup generator
  3. Verification: Run gates, generate, evaluate, human checkpoint
  4. Persistence: Save state, update ledger
  5. Scheduling: Determine next run

Deterministic Gates

Gates run BEFORE LLM calls to catch issues early:

fromloop_engineimportSyntaxGate, SecurityGate, GateRunnerrunner=GateRunner()
runner.add_gate(SyntaxGate())
runner.add_gate(SecurityGate())
result=runner.run_all(context)
# If any gate fails, we don't waste tokens on the LLM

Generator/Evaluator Separation

Prevents the "Ego Loop" where the LLM evaluates its own output:

# Generator: High temperature for creativitygenerator=GeneratorConfig(
model="claude-3-sonnet-20240229",
temperature=0.7,
system_prompt="You are a code generator..."
)
# Evaluator: Low temperature, skepticalevaluator=EvaluatorConfig(
model="claude-3-opus-20240229", # Different model!temperature=0.0, # Deterministicsystem_prompt="You are a skeptical code reviewer..."
)

Documentation


Real-World Stories

See stories/ for real-world use cases:

  • Stripe: Deterministic gates for payment processing
  • Anthropic: Evaluation infrastructure
  • OpenAI: Safety-critical systems
  • Your Story Here: Submit a story

Contributing

We welcome contributions! See CONTRIBUTING.md for:

  • Development setup
  • Code standards
  • PR process
  • Adding new patterns

Quick Development Setup

git clone https://github.com/chillum-codeX/loop-engineering.git
cd loop-engineering
pip install -e ".[dev]"
pytest tests/

License

MIT License - see LICENSE file.


Acknowledgments

  • Informed by public generator/evaluator and agent-loop engineering patterns
  • Inspired by Stripe's Minions pattern for deterministic gates
  • State machine patterns from classical control systems

Citation

If you use Loop Engineering in your research, please cite:

@software{loop_engineering,
title={Loop Engineering: A Framework for Autonomous AI Systems},
author={Loop Engineering Team},
year={2026},
url={https://github.com/chillum-codeX/loop-engineering}
}

DocumentationGitHub SetupQuickstart

About

A Python runtime for loop engineering with agent loops, budgets, verification, recovery, and persistence.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages