Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Gated runs v1: PreToolUse enforcement hook + policy engine - #3

Merged
githubscum merged 1 commit into
mainfrom
feat/gated-runs
Jul 22, 2026
Merged

Gated runs v1: PreToolUse enforcement hook + policy engine#3
githubscum merged 1 commit into
mainfrom
feat/gated-runs

Conversation

@githubscum

Copy link
Copy Markdown
Owner

Turns "gated runs" from a convention into a mode. Follows #1 (install readiness, session receipts) and #2 (subsessions).

Why a hook and not just the MCP tool

gated_action is a tool the agent chooses to call. Under bypassPermissions nothing stops a model from skipping it and calling Bash directly, which makes it a behavioral mitigation. PreToolUse sits below the agent's choice and can actually block (exit 2 denies the call; SessionEnd cannot). The two compose: the hook is the enforcement backstop, gated_action remains the request path.

The gate engine is reused, not rebuilt. verifyApproval and gatedAction in src/gate/ already do token binding, canonicalization, nonce replay protection and dual-direction receipts. The hook is an adapter over the same engine.

The policy

Locked before the build, so the enforcement point knew what it was enforcing. Sized to a budget of roughly 0-5 signings on a normal day, because a gate that gets turned off is worse than no gate.

RuleModeMatches
self-modgatesettings.json, gate/policy source, keys, policy.json
push-forcewarngit push --force
push-protectedwarnexplicit push to main/master
publishwarngh pr merge, npm publish, releases
egress-otherwarnPOST/PUT off-box, ssh/scp/rsync
destructivewarnrecursive delete outside a scratch allowlist
scope-escalationwarnscheduled tasks, service installs
spendoffdeferred: not reliably detectable from tool-call shape

Self-modification gates from day one because a gate that cannot protect itself is theater. Everything else starts in warn (receipt written, call allowed) for a calibration week, then promotes on the evidence. Feature-branch pushes, reads, in-repo edits, tests and local commits are unmatched and flow free.

Rule logic lives in code; <LOTOR_HOME>/policy.json only sets modes per rule. Editing that file is itself the self-mod rule.

Fail posture

  • Fail open on engine error. Unreadable policy, evaluator crash, store I/O failure: log to stderr, append a best-effort engine-error receipt, exit 0. A Lotor bug must not brick every tool call.
  • Fail closed at the token layer. Missing, invalid, or replayed token for a gate rule is a denial, receipted, exit 2. Invalid tokens are deleted so they cannot re-fire on every subsequent call.
  • Every denial prints the exact canonicalized request the owner must sign, plus the signing command.

Verification

  • 172 tests passing (96 baseline + 69 policy unit + 7 hook integration).
  • Falsification, as required: with token verification disabled, the replay test approves when it must deny. Restored, it denies. A gate test that passes with verification broken is worthless, so both runs were required and produced.
  • Manual CLI demo: warn rule exits 0 with a receipt, gate rule exits 2 with a denial receipt and signing instructions, unmatched call exits 0 with no chain I/O.

Honesty

KNOWN-LIMITS gains item 11: the gate matches patterns and patterns are evadable. v1's threat model is the honest agent making a consequential move without sign-off, not an adversarial model engineering around the matcher. Also documented: engine errors fail open by design, a bare git push to a checked-out protected branch is not detected, and hook registration lives in the user's settings file so a hostile edit between sessions is only caught at the next session's snapshot.

Activation

The hook is not registered by this PR. Registration is a PreToolUse block in the user's own settings.json, which is consistent with self-mod being rule 1.

Implements the locked policy of 2026-07-22: gate self-modification from
day one, warn everything else for a calibration week. Policy table, fail
posture, and approval flow per the spec; rule logic lives in code and
policy.json sets modes only.
- src/policy/: DEFAULT_POLICY, loadPolicy (writes the default on first
use, falls back without overwriting on malformed input), evaluate with
first-match-wins ordering, exported per-rule matchers.
- bin/hook-pre-tool-use.js: exit 2 blocks, exit 0 allows. Fail-open on
engine error (with an engine-error receipt for visibility), fail-closed
at the token layer. Tokens live in <home>/pending-approvals/; the first
valid one is consumed via gatedAction (nonce recorded exactly once),
invalid tokens are deleted so they cannot re-fire, and every denial
prints the exact canonicalized request the owner must sign.
- views: POLICY WARNINGS block in the morning-after summary.
- KNOWN-LIMITS 11: the gate matches patterns and patterns are evadable;
v1 threat model is the honest agent, not an adversarial one.
Tests: 172 passing (96 baseline + 69 policy unit + 7 hook integration).
Token path proven by falsification: with verification disabled the replay
test approves when it must deny; restored, all pass.
Note for the record: the executor ran one hook invocation against the
real store during its own verification (self-disclosed), leaving a true
policy-warn receipt at seq 2 and the default policy.json. The chain
verifies; the receipt stays, because it records something that actually
happened.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@githubscum