Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

GraphARC

GraphARC

PyPIDownloadsPythonCILicense

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc # Python >= 3.12
grapharc demo stage0 # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start # guided tour
grapharc init # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b # propose -> admit -> save
grapharc go # execute the saved plan
grapharc plan "..." --scripted # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs # live browser view of every run
grapharc replay <trace><run-id># reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

fromgrapharcimportGraphARC, GraphARCState, Budgetfromgrapharc.runtime.graphimportSTART, ENDclassState(GraphARCState):
question: stranswer: str=""defanswer(state: State) ->dict:
return {"answer": f"42 (asked: {state.question})"}
g=GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"}) # undeclared writes raiseg.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal : investigate the checkout outage
model : scripted stand-in (--scripted)
registry : grapharc.examples.plan_incident:build_registry
kinds : deploy, patch, triage, verify
policy : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow) [registry-default]
config : no grapharc.toml (flags and defaults only)
stopped : goal_met (the goal check was satisfied)
rounds : 2 of max 8
round 1: rejected nodes=2 executed=False rejected: edge_denied
round 2: admitted nodes=3 executed=True
state : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

frompydanticimportBaseModelfromgrapharc.harness.permissionsimportDecisionfromgrapharc.plannerimport (
AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
fromgrapharc.runtime.budgetimportBudgetfromgrapharc.testingimportScriptedChatModelclassState(BaseModel):
found: str=""fixed: str=""deffactory(spec): # bodies come from HERE, never a proposaldefbody(state):
return {"found": "cause"} ifspec.name=="search"else {"fixed": "patch"}
returnbodyregistry=NodeRegistry([ # the kinds a planner may proposeNodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
NodeSpec(name="edit", factory=factory, worst_case=CostEstimate(tokens=2000)),
NodeSpec(name="deploy", factory=factory),
]).freeze()
policy=EdgePolicy(rules=( # deny -> ask -> allow, unmatched is denyEdgeRule(action=Decision.DENY, target="deploy"),
EdgeRule(action=Decision.ALLOW),
))
plan='{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'loop=GovernedLoop(
planner=PlannerNode(
ScriptedChatModel(responses=[plan% ("deploy", "deploy"), plan% ("edit", "edit")]),
catalog=registry.catalog(),
),
checker=AdmissionChecker(registry=registry, edge_policy=policy),
materializer=Materializer(
registry=registry, state_schema=State,
writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
),
budget=Budget(max_tokens=100_000),
limits=LoopLimits(max_rounds=8),
goal_reached=lambdas: bool(s.fixed),
)
result=loop.run("find and fix the bug", State())
print(result.stop.value)
forrecordinresult.rounds:
print(record.round, record.admission.status.value, record.executed)
print([r.codeforrinresult.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
--model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARCClaude CodeOpenClawraw LangGraph
ShapeGoverned multi-node graph runtimeInteractive single-agent loopPersonal AI assistant gatewayGraph mechanism library
Who authorizes workDeterministic gate, pre-execution, with reasonsA human, live, per actionConfiguration and allowlistsNobody — convention
Cost controlWorst-case admission + per-node bill, fail-closedUsage visibilitySpend settingsNone built in
AuditOne replayable JSONL traceSession transcriptsLogsCheckpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.7 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages