[Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

Description

@eddy-curly

Area

apps/server

Problem or use case

An overnight agent orchestration run stops silently and stays stopped until a human notices.

Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

Stall AStall B
Turn marked completed00:19:19Z07:40:44Z
Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
Then total silence for3h 31m6h 50m
Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

00:10:20Z Did you get started?
00:11:05Z Any updates?
00:13:53Z What happening here? Are you still working on the overnight tasks?
00:18:55Z So you are working on it now? No need to schedule for later
04:23:48Z Can you continue the process in a /loop?
05:49:31Z Continuing?
06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
06:42:14Z Are you working through it?
07:15:17Z Continuing right?
14:39:06Z You stopped again
14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over

Proposed solution

1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

Why this matters

Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

Smallest useful scope

  • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
  • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

Cron-scheduled thread creation is explicitly out of scope for a first pass.

Alternatives considered

Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

Risks or tradeoffs

  • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
  • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
  • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

Examples or references

The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

      Description

      @eddy-curly

      Area

      apps/server

      Problem or use case

      An overnight agent orchestration run stops silently and stays stopped until a human notices.

      Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

      Stall AStall B
      Turn marked completed00:19:19Z07:40:44Z
      Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
      Then total silence for3h 31m6h 50m
      Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

      Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

      In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

      Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

      All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

      00:10:20Z Did you get started?
      00:11:05Z Any updates?
      00:13:53Z What happening here? Are you still working on the overnight tasks?
      00:18:55Z So you are working on it now? No need to schedule for later
      04:23:48Z Can you continue the process in a /loop?
      05:49:31Z Continuing?
      06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
      06:42:14Z Are you working through it?
      07:15:17Z Continuing right?
      14:39:06Z You stopped again
      14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
      

      Proposed solution

      1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

      A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

      2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

      Why this matters

      Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

      Smallest useful scope

      • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
      • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

      Cron-scheduled thread creation is explicitly out of scope for a first pass.

      Alternatives considered

      Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

      Risks or tradeoffs

      • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
      • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
      • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

      Examples or references

      The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

          Description

          @eddy-curly

          Area

          apps/server

          Problem or use case

          An overnight agent orchestration run stops silently and stays stopped until a human notices.

          Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

          Stall AStall B
          Turn marked completed00:19:19Z07:40:44Z
          Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
          Then total silence for3h 31m6h 50m
          Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

          Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

          In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

          Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

          All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

          00:10:20Z Did you get started?
          00:11:05Z Any updates?
          00:13:53Z What happening here? Are you still working on the overnight tasks?
          00:18:55Z So you are working on it now? No need to schedule for later
          04:23:48Z Can you continue the process in a /loop?
          05:49:31Z Continuing?
          06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
          06:42:14Z Are you working through it?
          07:15:17Z Continuing right?
          14:39:06Z You stopped again
          14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
          

          Proposed solution

          1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

          A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

          2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

          Why this matters

          Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

          Smallest useful scope

          • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
          • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

          Cron-scheduled thread creation is explicitly out of scope for a first pass.

          Alternatives considered

          Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

          Risks or tradeoffs

          • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
          • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
          • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

          Examples or references

          The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

              Description

              @eddy-curly

              Area

              apps/server

              Problem or use case

              An overnight agent orchestration run stops silently and stays stopped until a human notices.

              Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

              Stall AStall B
              Turn marked completed00:19:19Z07:40:44Z
              Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
              Then total silence for3h 31m6h 50m
              Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

              Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

              In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

              Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

              All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

              00:10:20Z Did you get started?
              00:11:05Z Any updates?
              00:13:53Z What happening here? Are you still working on the overnight tasks?
              00:18:55Z So you are working on it now? No need to schedule for later
              04:23:48Z Can you continue the process in a /loop?
              05:49:31Z Continuing?
              06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
              06:42:14Z Are you working through it?
              07:15:17Z Continuing right?
              14:39:06Z You stopped again
              14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
              

              Proposed solution

              1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

              A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

              2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

              Why this matters

              Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

              Smallest useful scope

              • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
              • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

              Cron-scheduled thread creation is explicitly out of scope for a first pass.

              Alternatives considered

              Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

              Risks or tradeoffs

              • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
              • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
              • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

              Examples or references

              The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

                  Description

                  @eddy-curly

                  Area

                  apps/server

                  Problem or use case

                  An overnight agent orchestration run stops silently and stays stopped until a human notices.

                  Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

                  Stall AStall B
                  Turn marked completed00:19:19Z07:40:44Z
                  Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
                  Then total silence for3h 31m6h 50m
                  Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

                  Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

                  In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

                  Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

                  All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

                  00:10:20Z Did you get started?
                  00:11:05Z Any updates?
                  00:13:53Z What happening here? Are you still working on the overnight tasks?
                  00:18:55Z So you are working on it now? No need to schedule for later
                  04:23:48Z Can you continue the process in a /loop?
                  05:49:31Z Continuing?
                  06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
                  06:42:14Z Are you working through it?
                  07:15:17Z Continuing right?
                  14:39:06Z You stopped again
                  14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
                  

                  Proposed solution

                  1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

                  A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

                  2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

                  Why this matters

                  Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

                  Smallest useful scope

                  • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
                  • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

                  Cron-scheduled thread creation is explicitly out of scope for a first pass.

                  Alternatives considered

                  Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

                  Risks or tradeoffs

                  • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
                  • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
                  • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

                  Examples or references

                  The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

                      Description

                      @eddy-curly

                      Area

                      apps/server

                      Problem or use case

                      An overnight agent orchestration run stops silently and stays stopped until a human notices.

                      Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

                      Stall AStall B
                      Turn marked completed00:19:19Z07:40:44Z
                      Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
                      Then total silence for3h 31m6h 50m
                      Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

                      Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

                      In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

                      Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

                      All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

                      00:10:20Z Did you get started?
                      00:11:05Z Any updates?
                      00:13:53Z What happening here? Are you still working on the overnight tasks?
                      00:18:55Z So you are working on it now? No need to schedule for later
                      04:23:48Z Can you continue the process in a /loop?
                      05:49:31Z Continuing?
                      06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
                      06:42:14Z Are you working through it?
                      07:15:17Z Continuing right?
                      14:39:06Z You stopped again
                      14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
                      

                      Proposed solution

                      1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

                      A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

                      2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

                      Why this matters

                      Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

                      Smallest useful scope

                      • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
                      • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

                      Cron-scheduled thread creation is explicitly out of scope for a first pass.

                      Alternatives considered

                      Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

                      Risks or tradeoffs

                      • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
                      • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
                      • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

                      Examples or references

                      The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

                          Description

                          @eddy-curly

                          Area

                          apps/server

                          Problem or use case

                          An overnight agent orchestration run stops silently and stays stopped until a human notices.

                          Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

                          Stall AStall B
                          Turn marked completed00:19:19Z07:40:44Z
                          Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
                          Then total silence for3h 31m6h 50m
                          Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

                          Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

                          In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

                          Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

                          All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

                          00:10:20Z Did you get started?
                          00:11:05Z Any updates?
                          00:13:53Z What happening here? Are you still working on the overnight tasks?
                          00:18:55Z So you are working on it now? No need to schedule for later
                          04:23:48Z Can you continue the process in a /loop?
                          05:49:31Z Continuing?
                          06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
                          06:42:14Z Are you working through it?
                          07:15:17Z Continuing right?
                          14:39:06Z You stopped again
                          14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
                          

                          Proposed solution

                          1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

                          A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

                          2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

                          Why this matters

                          Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

                          Smallest useful scope

                          • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
                          • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

                          Cron-scheduled thread creation is explicitly out of scope for a first pass.

                          Alternatives considered

                          Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

                          Risks or tradeoffs

                          • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
                          • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
                          • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

                          Examples or references

                          The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              [Feature]: supervise long-running threads — turns complete while background subagents are still working, and nothing restarts the run #38

                              Description

                              @eddy-curly

                              Area

                              apps/server

                              Problem or use case

                              An overnight agent orchestration run stops silently and stays stopped until a human notices.

                              Evidence from a real run — thread 3a85bdd3-8264-474c-aabc-a1a62eaf5d3a "Overnight TDD agent orchestration" (project eddy-knowledge-hub), 31 turns, 2026-07-30T21:34Z → 2026-07-31T14:54Z:

                              Stall AStall B
                              Turn marked completed00:19:19Z07:40:44Z
                              Background activities still arriving afterward1,358 (until 00:52:43Z, +33m)61 (until 07:48:47Z, +8m)
                              Then total silence for3h 31m6h 50m
                              Ended byhuman: "Can you continue the process in a /loop?"human: "You stopped again"

                              Stall A's 1,358 orphaned activities break down as 715 task.progress, 57 task.started, 58 task.completed, 527 context-window.updated, 1 checkpoint.captured. In other words, subagents kept working for 33 minutes inside a turn the server had already closed.

                              In Stall B a local_workflow task started activity fires at 07:40:32Z — 12 seconds before the turn closes at 07:40:44Z.

                              Both turns report state: completed with no error tone anywhere in the thread. The thread is not settled, snoozed, archived, or rate-limited: it has zero t3x.auto-resume.* activities and no entry in t3x-auto-resume.json, so auto-resume was never involved and could not have recovered it.

                              All 18 user messages in the thread are human (no message id carries the t3x-auto-resume: prefix). There is no automated continuation anywhere — the human is the loop, all night:

                              00:10:20Z Did you get started?
                              00:11:05Z Any updates?
                              00:13:53Z What happening here? Are you still working on the overnight tasks?
                              00:18:55Z So you are working on it now? No need to schedule for later
                              04:23:48Z Can you continue the process in a /loop?
                              05:49:31Z Continuing?
                              06:00:20Z Great. Continue working. I am going to bed now, ensure that your sub agents are not dying
                              06:42:14Z Are you working through it?
                              07:15:17Z Continuing right?
                              14:39:06Z You stopped again
                              14:53:19Z Can you note down your current progress in PROGRESS.md? I will have another agent take over
                              

                              Proposed solution

                              1. Supervision. Let a thread be marked as a long-running/supervised run. When its turn completes but the run is not actually done, T3 Code re-prompts it instead of leaving it idle.

                              A stall signal is already derivable from existing projections: activities (task.progress, task.started, task.completed) continuing to arrive after projection_turns.completed_at, then going quiet with no new turn requested.

                              2. Make it visible and pinned. A distinct thread type that pins to the top of the sidebar and is visually differentiated, showing live state (running / stalled / last activity age) so an overnight run is identifiable at a glance rather than sinking into the list.

                              Why this matters

                              Unattended overnight work is the entire point of leaving an orchestration running. Today a stall costs the whole remaining window — 6h50m in the case above — and is only discoverable by manually opening the thread and reading its timeline. The user cannot tell a finished run from a dead one without reading it.

                              Smallest useful scope

                              • A per-thread "supervised" flag, plus pinned/distinct rendering in the web sidebar.
                              • A server-side stall detector that re-prompts once when a supervised thread has been idle for N minutes with no active turn.

                              Cron-scheduled thread creation is explicitly out of scope for a first pass.

                              Alternatives considered

                              Relying on the agent's own /loop self-scheduling. This is what was tried and it is what failed — the agent's wakeups do not survive the turn boundary, so once the turn closes there is nothing left to re-invoke it. Manual nudging works but requires a human awake at 4am.

                              Risks or tradeoffs

                              • Re-prompting a genuinely finished run wastes tokens and can confuse the agent. Needs a done-condition and an attempt cap — auto-resume's maxResumesPer24h is prior art worth reusing rather than reinventing.
                              • Interaction with settled / snoozed and with auto-resume's guards needs an explicit decision, not two mechanisms racing to revive the same thread.
                              • Per AGENTS.md: needs a decision for web, desktop, and mobile sidebars, and a reverse state (unpin / unsupervise) — a one-way door is a bug.

                              Examples or references

                              The mechanism described here — a turn reporting completed while its background tasks are still streaming activities — may deserve its own bug report. It is documented here because it is the direct cause of the stalls, and because a supervisor needs to know about it either way.

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions