A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

Description

@os-trump

Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

The failure, measured end to end

case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
Node 'notify_team' failed: notify: at least one recipient is required,
but every recipient template resolved to nothing: {currentCase.owner_id}

crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

run statussummaryflaggednotified
beforefailed{selected: 5, acted: 0}2 of 51

The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

Why this is the platform's to answer, not the app's

The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

  • any node that can fail on one row takes down a sweep over all rows;
  • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
  • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

The ask

A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

Shape suggestion only; the naming is yours:

{id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

Prior art checked before filing

Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

Downstream reference

objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

Generated by Claude Code

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

      Description

      @os-trump

      Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

      The failure, measured end to end

      case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

      ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
      Node 'notify_team' failed: notify: at least one recipient is required,
      but every recipient template resolved to nothing: {currentCase.owner_id}
      

      crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

      builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

      Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

      run statussummaryflaggednotified
      beforefailed{selected: 5, acted: 0}2 of 51

      The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

      Why this is the platform's to answer, not the app's

      The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

      • any node that can fail on one row takes down a sweep over all rows;
      • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
      • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

      An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

      The ask

      A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

      Shape suggestion only; the naming is yours:

      {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

      with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

      Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

      Prior art checked before filing

      Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

      Downstream reference

      objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

      Generated by Claude Code

      Metadata

      Metadata

      Assignees

      No one assigned

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

          Description

          @os-trump

          Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

          The failure, measured end to end

          case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

          ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
          Node 'notify_team' failed: notify: at least one recipient is required,
          but every recipient template resolved to nothing: {currentCase.owner_id}
          

          crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

          builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

          Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

          run statussummaryflaggednotified
          beforefailed{selected: 5, acted: 0}2 of 51

          The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

          Why this is the platform's to answer, not the app's

          The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

          • any node that can fail on one row takes down a sweep over all rows;
          • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
          • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

          An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

          The ask

          A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

          Shape suggestion only; the naming is yours:

          {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

          with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

          Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

          Prior art checked before filing

          Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

          Downstream reference

          objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

          Generated by Claude Code

          Metadata

          Metadata

          Assignees

          No one assigned

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

              Description

              @os-trump

              Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

              The failure, measured end to end

              case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

              ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
              Node 'notify_team' failed: notify: at least one recipient is required,
              but every recipient template resolved to nothing: {currentCase.owner_id}
              

              crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

              builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

              Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

              run statussummaryflaggednotified
              beforefailed{selected: 5, acted: 0}2 of 51

              The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

              Why this is the platform's to answer, not the app's

              The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

              • any node that can fail on one row takes down a sweep over all rows;
              • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
              • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

              An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

              The ask

              A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

              Shape suggestion only; the naming is yours:

              {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

              with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

              Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

              Prior art checked before filing

              Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

              Downstream reference

              objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

              Generated by Claude Code

              Metadata

              Metadata

              Assignees

              No one assigned

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

                  Description

                  @os-trump

                  Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

                  The failure, measured end to end

                  case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

                  ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
                  Node 'notify_team' failed: notify: at least one recipient is required,
                  but every recipient template resolved to nothing: {currentCase.owner_id}
                  

                  crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

                  builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

                  Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

                  run statussummaryflaggednotified
                  beforefailed{selected: 5, acted: 0}2 of 51

                  The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

                  Why this is the platform's to answer, not the app's

                  The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

                  • any node that can fail on one row takes down a sweep over all rows;
                  • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
                  • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

                  An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

                  The ask

                  A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

                  Shape suggestion only; the naming is yours:

                  {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

                  with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

                  Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

                  Prior art checked before filing

                  Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

                  Downstream reference

                  objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

                  Generated by Claude Code

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

                      Description

                      @os-trump

                      Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

                      The failure, measured end to end

                      case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

                      ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
                      Node 'notify_team' failed: notify: at least one recipient is required,
                      but every recipient template resolved to nothing: {currentCase.owner_id}
                      

                      crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

                      builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

                      Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

                      run statussummaryflaggednotified
                      beforefailed{selected: 5, acted: 0}2 of 51

                      The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

                      Why this is the platform's to answer, not the app's

                      The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

                      • any node that can fail on one row takes down a sweep over all rows;
                      • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
                      • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

                      An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

                      The ask

                      A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

                      Shape suggestion only; the naming is yours:

                      {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

                      with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

                      Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

                      Prior art checked before filing

                      Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

                      Downstream reference

                      objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

                      Generated by Claude Code

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

                          Description

                          @os-trump

                          Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

                          The failure, measured end to end

                          case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

                          ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
                          Node 'notify_team' failed: notify: at least one recipient is required,
                          but every recipient template resolved to nothing: {currentCase.owner_id}
                          

                          crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

                          builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

                          Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

                          run statussummaryflaggednotified
                          beforefailed{selected: 5, acted: 0}2 of 51

                          The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

                          Why this is the platform's to answer, not the app's

                          The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

                          • any node that can fail on one row takes down a sweep over all rows;
                          • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
                          • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

                          An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

                          The ask

                          A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

                          Shape suggestion only; the naming is yours:

                          {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

                          with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

                          Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

                          Prior art checked before filing

                          Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

                          Downstream reference

                          objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

                          Generated by Claude Code

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              A loop node aborts the entire flow run when one iteration's node fails — a single bad row kills a whole scheduled sweep, and there is no per-iteration containment to opt into #13681

                              Description

                              @os-trump

                              Platform ask from objectstack-ai/hotcrm, the reference CRM app. Measured on @objectstack/runtime17.1.0. Nothing is mis-implemented against a stated contract — there is no contract here to state, which is the ask.

                              The failure, measured end to end

                              case_sla_monitor is a scheduled sweep: select every breached case, then for each one flag the breach and notify the owner. On a fresh boot it died terminally:

                              ERROR Trigger-fired run of flow 'case_sla_monitor' failed (trigger 'schedule')
                              Node 'notify_team' failed: notify: at least one recipient is required,
                              but every recipient template resolved to nothing: {currentCase.owner_id}
                              

                              crm_case.owner_id is nullable and an ownerless case is an ordinary state. The path is:

                              builtin notify returns success: falseexecuteNode throws → runRegion rethrows → the loop node awaits it with no try/catch → the run dies.

                              Reproduced deterministically in the app's own harness against the real AutomationEngine (five breached cases, the ownerless one third in line):

                              run statussummaryflaggednotified
                              beforefailed{selected: 5, acted: 0}2 of 51

                              The two cases after the ownerless one were never processed at all. One row with a null lookup cost the other 60% of the sweep, and nothing retried it.

                              Why this is the platform's to answer, not the app's

                              The app has fixed its instance by inserting a decision gateway so notify is never reached with an empty audience. That works and it ships. But it is a per-call-site guard, and the general shape it patches is not specific to notify:

                              • any node that can fail on one row takes down a sweep over all rows;
                              • the blast radius is invisible at author time — nothing in objectstack validate, build or the flow lint family says "this loop has no containment";
                              • the damage is silent in the only way that matters: the other rows produce no error of their own. They simply never run.

                              An app can only defend by hand-writing a predicate in front of every fallible node inside every loop body, for every failure mode that node has. That is not a contract, it is a memory test — and the app's own first draft of exactly such a predicate was wrong (a CEL != '' test is true for ' ', while notify trims before dropping), which is a small demonstration of the point.

                              The ask

                              A declared, opt-in per-iteration failure containment on the loop node — continue to the next iteration, record the failure against that iteration, and report it in the run summary rather than aborting the run.

                              Shape suggestion only; the naming is yours:

                              {id: 'loop_cases',type: 'loop',config: {onIterationError: 'continue'|'abort'}}

                              with 'abort' the default so nothing changes for existing flows. The run summary already carries selected / acted / skipped (#4354), so failed-iteration counts have a natural home beside them, and a run that partially succeeded stops being indistinguishable from one that died at row 1.

                              Whether the default should eventually flip is a separate question — for a scheduled sweep "process the rest and tell me what failed" is almost always the intended semantics, and "abort" is the one an author is least likely to have chosen deliberately.

                              Prior art checked before filing

                              Searched the flow/loop neighbourhood: #5633, #5383 and #4347 are all about lint and conversion passes failing to descend into loop bodies; #4354 surfaced run summaries; #3712 and #3427 are unrelated trigger/context issues. None covers runtime failure containment inside the loop. If this is a duplicate of something I could not surface, close it against that card.

                              Downstream reference

                              objectstack-ai/hotcrm#1405 carries the full reproduction and the per-site workaround the app shipped. A sibling ask covering the narrower half — notify treating an empty audience as a hard failure rather than a recorded skip — is filed separately; this card is the general one and would retire the need for the guard at every call site, not just notify's.

                              Generated by Claude Code

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions