Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -7,3 +7,6 @@ __pycache__/
dist/
build/
*.log
.worktrees/
docs/superpowers/
new-prd.md
141 changes: 121 additions & 20 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -21,6 +21,19 @@

**Evolution Kernel** is a minimal protocol and runtime for autonomous, self-evolving software systems.

## Quick Start

```bash
# Install
pip install -e .

# Run the demo (uses fixture roles from tests/fixtures/)
bash examples/run_demo.sh

# Check the result
cat /tmp/ek-demo-ledger/runs/0001/decision.json
```

It is not a project-specific automation script. Its purpose is to make software evolution **controlled, reproducible, sandboxed, auditable, and reversible**. Any project can become an optimization target once it can expose a goal, a sandbox, and an evaluator.

## Why It Exists
Expand All@@ -39,11 +52,15 @@ Evolution Kernel provides that loop as a small, inspectable runtime.

```mermaid
flowchart LR
Goal[Goal] --> Governor[Governor]
Governor --> Planner[Planner]
Config[Config] --> Governor[Governor]
Governor --> Observer[Observer]
Observer --> Obs[observation.json]
Obs --> Planner[Planner]
Planner --> Plan[plan.json]
Plan --> Executor[Executor]
Executor --> Candidate[Sandbox candidate]
Executor --> Scope{Scope check}
Scope -- violation --> Reject[Reject + ledger]
Scope -- ok --> Candidate[Sandbox candidate]
Candidate --> Evaluator[Evaluator]
Evaluator --> Eval[evaluation.json]
Eval --> Governor
Expand All@@ -59,31 +76,65 @@ Token-Ignition is therefore the first optimization target and reference adapter,

## Current Status

The current v0 implementation provides the foundational runtime:

| Area | What exists now |
| --- | --- |
| Governor | Deterministic orchestration for planning, execution, evaluation, promotion, rollback, and ledger updates. |
| Sandbox | Git worktree-based experiment isolation. Candidate changes do not affect the accepted branch unless promoted. |
| Observer | Collects evidence from local files and shell commands before planning; writes `observation.json` into the ledger. |
| Mutation scope | `allowed_paths` enforces which files the executor may touch; violations are recorded as `scope_violation` without calling the evaluator. |
| Hard stops | `max_iterations` and `max_consecutive_failures` limits persist across runs in `ledger/state.json`; `--reset` clears them. |
| YAML config | `evolution.yml` unifies mission, evidence sources, mutation scope, hard stops, and role commands in one file. |
| Role handoff | `planner`, `executor`, and `evaluator` run as isolated commands and communicate through JSON files. |
| Promotion model | Accepted candidates advance the local `evolution/accepted` branch. Rejected experiments remain recorded but do not advance it. |
| First adapter | A Token-Ignition adapter with a hand-written golden set for evaluator evolution. |

## What It Does Not Do Yet
## Acceptance Checklist

Run the following scenarios to verify the kernel works end-to-end:

```bash
# 1. Full happy path (accept)
bash examples/run_demo.sh
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: true

# 2. Evaluator rejects → decision recorded, no promotion
# Edit examples/run_demo.sh to use evaluator_reject.py, re-run, check:
cat /tmp/ek-demo-ledger/runs/0001/decision.json # expect accepted: false

# 3. Observer writes evidence
cat /tmp/ek-demo-ledger/runs/0001/observation.json # expect sources[] with results

# 4. Mutation scope violation → scope_violation, evaluator NOT called
# (see test_scope_violation_rejects_without_calling_evaluator in tests/test_governor.py)

# 5. Hard stop (max_consecutive_failures)
bash examples/run_demo_hard_stop.sh # expect HardStopError after 2 failures

| Not yet | Why it matters |
# 6. Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/ek-demo-ledger
# Re-run → expect normal execution again
```

Or run the full unit test suite:

```bash
python3 -m unittest discover -s tests -v
```

## Current Limitations

| Limitation | Detail |
| --- | --- |
| LLM-native planner/executor | The current tests use fixture scripts; real agent integrations are the next step. |
| Strong process/container sandboxing | Git worktrees isolate files, but executor and evaluator isolation should become stronger. |
| Multi-target adapter framework | Token-Ignition is the first target; more adapters are needed to prove generality. |
| Parallel evolution branches | v0 focuses on one accepted branch and a simple promotion path. |
| No LLM-native roles | Planner/executor/evaluator are shell scripts or Python fixtures; real agent integrations are the next step. |
| File-only sandbox | Git worktrees isolate files; process/container-level isolation is not yet enforced. |
| Single evolution branch | v0 supports one `evolution/accepted` branch; parallel branches are not yet supported. |
| Scope check is path-prefix only | `allowed_paths` are matched by string prefix against `git status` output; glob or regex patterns are not supported. |

## Roadmap

- [ ] Add LLM-driven planner and executor implementations.
- [ ] Add stronger sandbox isolation for executor and evaluator runs.
- [ ] Strengthen sandbox isolation (process/container level).
- [ ] Generalize the adapter interface beyond Token-Ignition.
- [ ] Add examples for multiple project types.
- [ ] Support glob/regex patterns in `mutation_scope.allowed_paths`.
- [ ] Support parallel evolution branches and richer merge strategies.
- [ ] Improve reporting around ledger history, promotion decisions, and rejected candidates.

Expand All@@ -99,16 +150,26 @@ python3 -m unittest discover -s tests -v
python3 adapters/token_ignition/evaluate_golden_cases.py
```

## CLI Shape
## CLI

```bash
# Using a config file (recommended)
python3 -m evolution_kernel.cli \
--config examples/evolution.yml \
--repo /path/to/target-repo \
--ledger /path/to/evolution-ledger \
--goal /path/to/goal.json \
--planner python3 /path/to/planner.py \
--executor python3 /path/to/executor.py \
--evaluator python3 /path/to/evaluator.py
--ledger /tmp/evolution-ledger

# Legacy: explicit role args (backward compatible)
python3 -m evolution_kernel.cli \
--repo /path/to/target-repo \
--ledger /tmp/evolution-ledger \
--goal goal.json \
--planner python3 my_planner.py \
--executor python3 my_executor.py \
--evaluator python3 my_evaluator.py

# Reset hard stop state
python3 -m evolution_kernel.cli --reset --ledger /tmp/evolution-ledger
```

Each role command receives:
Expand All@@ -118,3 +179,43 @@ Each role command receives:
--output <json>
--worktree <sandbox path>
```

## Config format (`evolution.yml`)

```yaml
mission: "Improve the project."

evidence_sources:
- type: file
path: "./metrics.json"
- type: shell
command: "bash ./scripts/status.sh"

mutation_scope:
allowed_paths:
- "src/"
- "tests/"

hard_stops:
max_iterations: 3
max_consecutive_failures: 2

roles:
planner: ["python3", "my_planner.py"]
executor: ["python3", "my_executor.py"]
evaluator: ["python3", "my_evaluator.py"]
```

## Ledger artifacts (per run)

Each run produces the following files in `ledger/runs/{run_id}/`:

| File | Description |
|---|---|
| `observation.json` | Evidence collected before planning |
| `plan.json` | Planner output |
| `patch.diff` | Changes made by executor |
| `candidate_commit.txt` | SHA of the candidate commit |
| `evaluation.json` | Evaluator verdict |
| `decision.json` | Accept/reject decision with reason |
| `reflection.json` | Summary for auditing |
Loading