session: nothing reclaims a session whose runner died holding it #19

Description

@Shashankss1205

Summary

A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

That is the whole of the "one runner at a time" guard — it stops a second
runner from *claiming* a session, and it does not detect a runner that died
holding one. A session stuck in `running` after a crash has to be released
deliberately, which is a legal `running -> idle` transition.

But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

status after crash: running
runner_pid: 24442 alive: no (dead)
resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']

Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

Why this matters

The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

Where in the code

  • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
  • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
  • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
  • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
  • README.md:499 — the documented limitation
  • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

Confirm it:

uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

What to change

The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

  1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
  2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
  3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
  4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

How to verify

uv run pytest tests/test_session.py -q
uv run pytest -q
uv run ruff check .

The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

Acceptance criteria

  • A session whose runner died can be released without touching SQLite by hand
  • The release is refused while the recorded runner is alive
  • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
  • Open approval holds survive the reclaim; nothing runs unapproved because of it
  • A second runner cannot claim a genuinely running session (existing tests unchanged)
  • store.py, runtime.py docstrings and README.md:499 updated to match
  • uv run pytest stays green and uv run ruff check . is clean
  • Any README or cookbook sentence this changes is updated in the same pull request

Skill level — experience required

This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      session: nothing reclaims a session whose runner died holding it #19

      Description

      @Shashankss1205

      Summary

      A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

      That is the whole of the "one runner at a time" guard — it stops a second
      runner from *claiming* a session, and it does not detect a runner that died
      holding one. A session stuck in `running` after a crash has to be released
      deliberately, which is a legal `running -> idle` transition.
      

      But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

      status after crash: running
      runner_pid: 24442 alive: no (dead)
      resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
      

      Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

      README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

      Why this matters

      The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

      Where in the code

      • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
      • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
      • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
      • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
      • README.md:499 — the documented limitation
      • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

      Confirm it:

      uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

      What to change

      The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

      1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
      2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
      3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
      4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

      Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

      How to verify

      uv run pytest tests/test_session.py -q
      uv run pytest -q
      uv run ruff check .

      The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

      Acceptance criteria

      • A session whose runner died can be released without touching SQLite by hand
      • The release is refused while the recorded runner is alive
      • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
      • Open approval holds survive the reclaim; nothing runs unapproved because of it
      • A second runner cannot claim a genuinely running session (existing tests unchanged)
      • store.py, runtime.py docstrings and README.md:499 updated to match
      • uv run pytest stays green and uv run ruff check . is clean
      • Any README or cookbook sentence this changes is updated in the same pull request

      Skill level — experience required

      This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          session: nothing reclaims a session whose runner died holding it #19

          Description

          @Shashankss1205

          Summary

          A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

          That is the whole of the "one runner at a time" guard — it stops a second
          runner from *claiming* a session, and it does not detect a runner that died
          holding one. A session stuck in `running` after a crash has to be released
          deliberately, which is a legal `running -> idle` transition.
          

          But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

          status after crash: running
          runner_pid: 24442 alive: no (dead)
          resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
          

          Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

          README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

          Why this matters

          The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

          Where in the code

          • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
          • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
          • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
          • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
          • README.md:499 — the documented limitation
          • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

          Confirm it:

          uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

          What to change

          The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

          1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
          2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
          3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
          4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

          Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

          How to verify

          uv run pytest tests/test_session.py -q
          uv run pytest -q
          uv run ruff check .

          The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

          Acceptance criteria

          • A session whose runner died can be released without touching SQLite by hand
          • The release is refused while the recorded runner is alive
          • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
          • Open approval holds survive the reclaim; nothing runs unapproved because of it
          • A second runner cannot claim a genuinely running session (existing tests unchanged)
          • store.py, runtime.py docstrings and README.md:499 updated to match
          • uv run pytest stays green and uv run ruff check . is clean
          • Any README or cookbook sentence this changes is updated in the same pull request

          Skill level — experience required

          This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              session: nothing reclaims a session whose runner died holding it #19

              Description

              @Shashankss1205

              Summary

              A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

              That is the whole of the "one runner at a time" guard — it stops a second
              runner from *claiming* a session, and it does not detect a runner that died
              holding one. A session stuck in `running` after a crash has to be released
              deliberately, which is a legal `running -> idle` transition.
              

              But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

              status after crash: running
              runner_pid: 24442 alive: no (dead)
              resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
              

              Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

              README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

              Why this matters

              The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

              Where in the code

              • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
              • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
              • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
              • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
              • README.md:499 — the documented limitation
              • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

              Confirm it:

              uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

              What to change

              The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

              1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
              2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
              3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
              4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

              Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

              How to verify

              uv run pytest tests/test_session.py -q
              uv run pytest -q
              uv run ruff check .

              The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

              Acceptance criteria

              • A session whose runner died can be released without touching SQLite by hand
              • The release is refused while the recorded runner is alive
              • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
              • Open approval holds survive the reclaim; nothing runs unapproved because of it
              • A second runner cannot claim a genuinely running session (existing tests unchanged)
              • store.py, runtime.py docstrings and README.md:499 updated to match
              • uv run pytest stays green and uv run ruff check . is clean
              • Any README or cookbook sentence this changes is updated in the same pull request

              Skill level — experience required

              This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  session: nothing reclaims a session whose runner died holding it #19

                  Description

                  @Shashankss1205

                  Summary

                  A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

                  That is the whole of the "one runner at a time" guard — it stops a second
                  runner from *claiming* a session, and it does not detect a runner that died
                  holding one. A session stuck in `running` after a crash has to be released
                  deliberately, which is a legal `running -> idle` transition.
                  

                  But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

                  status after crash: running
                  runner_pid: 24442 alive: no (dead)
                  resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
                  

                  Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

                  README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

                  Why this matters

                  The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

                  Where in the code

                  • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
                  • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
                  • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
                  • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
                  • README.md:499 — the documented limitation
                  • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

                  Confirm it:

                  uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

                  What to change

                  The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

                  1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
                  2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
                  3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
                  4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

                  Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

                  How to verify

                  uv run pytest tests/test_session.py -q
                  uv run pytest -q
                  uv run ruff check .

                  The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

                  Acceptance criteria

                  • A session whose runner died can be released without touching SQLite by hand
                  • The release is refused while the recorded runner is alive
                  • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
                  • Open approval holds survive the reclaim; nothing runs unapproved because of it
                  • A second runner cannot claim a genuinely running session (existing tests unchanged)
                  • store.py, runtime.py docstrings and README.md:499 updated to match
                  • uv run pytest stays green and uv run ruff check . is clean
                  • Any README or cookbook sentence this changes is updated in the same pull request

                  Skill level — experience required

                  This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      session: nothing reclaims a session whose runner died holding it #19

                      Description

                      @Shashankss1205

                      Summary

                      A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

                      That is the whole of the "one runner at a time" guard — it stops a second
                      runner from *claiming* a session, and it does not detect a runner that died
                      holding one. A session stuck in `running` after a crash has to be released
                      deliberately, which is a legal `running -> idle` transition.
                      

                      But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

                      status after crash: running
                      runner_pid: 24442 alive: no (dead)
                      resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
                      

                      Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

                      README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

                      Why this matters

                      The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

                      Where in the code

                      • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
                      • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
                      • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
                      • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
                      • README.md:499 — the documented limitation
                      • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

                      Confirm it:

                      uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

                      What to change

                      The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

                      1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
                      2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
                      3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
                      4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

                      Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

                      How to verify

                      uv run pytest tests/test_session.py -q
                      uv run pytest -q
                      uv run ruff check .

                      The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

                      Acceptance criteria

                      • A session whose runner died can be released without touching SQLite by hand
                      • The release is refused while the recorded runner is alive
                      • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
                      • Open approval holds survive the reclaim; nothing runs unapproved because of it
                      • A second runner cannot claim a genuinely running session (existing tests unchanged)
                      • store.py, runtime.py docstrings and README.md:499 updated to match
                      • uv run pytest stays green and uv run ruff check . is clean
                      • Any README or cookbook sentence this changes is updated in the same pull request

                      Skill level — experience required

                      This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          session: nothing reclaims a session whose runner died holding it #19

                          Description

                          @Shashankss1205

                          Summary

                          A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

                          That is the whole of the "one runner at a time" guard — it stops a second
                          runner from *claiming* a session, and it does not detect a runner that died
                          holding one. A session stuck in `running` after a crash has to be released
                          deliberately, which is a legal `running -> idle` transition.
                          

                          But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

                          status after crash: running
                          runner_pid: 24442 alive: no (dead)
                          resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
                          

                          Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

                          README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

                          Why this matters

                          The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

                          Where in the code

                          • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
                          • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
                          • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
                          • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
                          • README.md:499 — the documented limitation
                          • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

                          Confirm it:

                          uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

                          What to change

                          The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

                          1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
                          2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
                          3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
                          4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

                          Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

                          How to verify

                          uv run pytest tests/test_session.py -q
                          uv run pytest -q
                          uv run ruff check .

                          The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

                          Acceptance criteria

                          • A session whose runner died can be released without touching SQLite by hand
                          • The release is refused while the recorded runner is alive
                          • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
                          • Open approval holds survive the reclaim; nothing runs unapproved because of it
                          • A second runner cannot claim a genuinely running session (existing tests unchanged)
                          • store.py, runtime.py docstrings and README.md:499 updated to match
                          • uv run pytest stays green and uv run ruff check . is clean
                          • Any README or cookbook sentence this changes is updated in the same pull request

                          Skill level — experience required

                          This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              session: nothing reclaims a session whose runner died holding it #19

                              Description

                              @Shashankss1205

                              Summary

                              A runner's claim on a session is a compare-and-set, not a lease. SessionStore.transition says so plainly:

                              That is the whole of the "one runner at a time" guard — it stops a second
                              runner from *claiming* a session, and it does not detect a runner that died
                              holding one. A session stuck in `running` after a crash has to be released
                              deliberately, which is a legal `running -> idle` transition.
                              

                              But no release path exists anywhere in the package: no method performs that "deliberate" transition, nothing consults runner_pid, and there is no CLI or API to do it either. Verified with a child process that claims the session and dies mid-turn (os._exit, standing in for a crash or OOM-kill):

                              status after crash: running
                              runner_pid: 24442 alive: no (dead)
                              resume attempt: SessionBusy - session 'crashme' is running; expected one of ['idle', 'interrupted', 'awaiting_approval', 'failed', 'created']
                              

                              Every later Session.run() raises SessionBusy forever. The record even carries the dead runner_pid, so the store knows who died holding it — and offers no way to act on that knowledge short of hand-editing SQLite.

                              README states the limitation ("a runner claim is a claim rather than a lease — nothing reclaims a session whose runner died holding it"), and #7 defers to "the separate lease issue" in its step 3. This is that issue.

                              Why this matters

                              The session layer's headline is durability: "a new process pointed at the same directory can pick a session up by id." A crash is precisely when that promise is needed, and it is precisely when it fails — the one process that could release the claim is the one that died. A session holding a human-approval gate wedges in running with its holds intact and its queued decisions unread; an operator's only remedy is sqlite3 sessions.sqlite "UPDATE sessions SET status=...", which bypasses the lifecycle validation and writes no transition row, corrupting the very audit trail the store exists to keep. Under #7 (HTTP workers restarting) this stops being an edge case and becomes routine.

                              Where in the code

                              • grapharc/session/store.py:393-399 — the transition() docstring quoted above: the claim guard, and the admission that nothing detects a dead runner
                              • grapharc/session/store.py:420-423runner_pid is stamped on the way into running and cleared on every way out, so a running row always names the process that would have to be dead
                              • grapharc/session/runtime.py:52-54 — the module docstring's honest bullet: "nothing reclaims one whose runner died holding it"
                              • grapharc/session/runtime.py:369-375 — the claim itself (expect=RESUMABLE), which is what every future runner loses against a wedged row
                              • README.md:499 — the documented limitation
                              • Issue server: the HTTP API does not use the durable session layer #7, "What to change" step 3 — names this as the separate lease issue

                              Confirm it:

                              uv run python - <<'EOF'import os, subprocess, sys, tempfile, textwrapfrom grapharc.session import SessionBusy, SessionManagerfrom grapharc.session.demo import GRAPH_NAMEroot = tempfile.mkdtemp()m = SessionManager(root)m.create(GRAPH_NAME, session_id="crashme")child = textwrap.dedent(f""" import os from grapharc.session import SessionStatus, SessionStore from grapharc.session.store import RESUMABLE store = SessionStore({os.path.join(root, 'sessions.sqlite')!r}) store.transition("crashme", SessionStatus.RUNNING, expect=RESUMABLE, reason="turn started") os._exit(1) # crash mid-turn""")subprocess.run([sys.executable, "-c", child])rec = m.store.require("crashme")print("status:", rec.status.value, "runner_pid:", rec.runner_pid)try: m.resume("crashme").run({"inbox": ["hello"]})except SessionBusy as exc: print("wedged forever:", exc)EOF

                              What to change

                              The hard part is semantic — a pid can be recycled, so "the pid is dead" is evidence and "the pid is alive" is not proof the runner is — which is why this proposes a deliberate, recorded release rather than an automatic one:

                              1. Add SessionStore.release_dead_runner(session_id, *, reason="") (name negotiable): under the existing BEGIN IMMEDIATE, verify the row is running, verify runner_pid names a process that no longer exists on this host (os.kill(pid, 0)ProcessLookupError; refuse when the pid is alive or is our own), then perform a legal running -> failed transition with last_error and the transition reason naming the dead pid. failed rather than idle, because a turn that died is a turn that did not settle — and failed is already re-runnable while keeping pending_approval intact, so an open hold survives the reclaim (the same reason _settle keeps holds on a failed turn).
                              2. Surface it on SessionManager (e.g. manager.reclaim(session_id)) so an operator does not have to touch the store class directly.
                              3. Decide whether Session.run() ever calls it automatically. Recommendation: not by default — pid liveness is host-local and pid reuse makes auto-reclaim a way for two live runners to fight — but this deserves a design comment before code.
                              4. Update the honesty paragraphs this obsoletes: store.py:393-399, runtime.py:52-54, and README.md:499 ("nothing reclaims" becomes "reclaimed deliberately via ...").

                              Deliberately out of scope: a heartbeat/expiry lease that works across hosts (the store is a local SQLite file; design that under #7 if the HTTP layer ever needs it), and any automatic background sweeper.

                              How to verify

                              uv run pytest tests/test_session.py -q
                              uv run pytest -q
                              uv run ruff check .

                              The decisive new test is cross-process, using the existing run_child helpers in tests/test_session.py: a child claims the session and os._exits; the parent reclaims, sees a running -> failed transition row naming the dead pid, and then runs the session to completion with every node appearing exactly once. A second test asserts reclaim refuses when the recorded runner is alive (a child parked on a barrier). Revert the source change and watch both go red.

                              Acceptance criteria

                              • A session whose runner died can be released without touching SQLite by hand
                              • The release is refused while the recorded runner is alive
                              • The release writes a transition row with the reason and the dead pid — the audit trail shows the reclaim rather than hiding it
                              • Open approval holds survive the reclaim; nothing runs unapproved because of it
                              • A second runner cannot claim a genuinely running session (existing tests unchanged)
                              • store.py, runtime.py docstrings and README.md:499 updated to match
                              • uv run pytest stays green and uv run ruff check . is clean
                              • Any README or cookbook sentence this changes is updated in the same pull request

                              Skill level — experience required

                              This spans the store's transaction discipline, the session lifecycle table, and cross-process crash semantics, and the failure modes are the quiet kind: pid reuse making a dead runner look alive, two processes reclaiming at once (the BEGIN IMMEDIATE pattern must carry the liveness check too), and a reclaim that accidentally drops an open hold. Please open a design comment here before writing code — in particular on whether reclaim lands on failed vs interrupted, and on whether run() may ever auto-reclaim. If you have built job-queue or workflow-lease systems, this is a well-shaped one to take.

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                enhancementNew feature or requestexperience requiredDeep familiarity with the codebase or domain needed; not a starter taskhelp wantedExtra attention is needed

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions