os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

Description

@os-trump

Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

What

Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

1192 uncaughtException code=EPIPE msg=write EPIPE
stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
| at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
1192 exit code=1

12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

Why it matters

Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

parent's readerexitelapsed
drained2~8.3 s (145696 bytes delivered)
paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
read end destroyed11.4 s

So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

The second consequence: writeStderr() is never reached

bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

bin/run-dev.js stateexitelapsed
pristine11387-1711 ms
write callback removed, so the closed path can only finish on the bound11517-1633 ms
process.stderr.on('error', noop) added, callback kept28787-8979 ms
both223601-23712 ms (one bound paid)

Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

What a fix owes

  1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
  2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
  3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

Related

Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

      Description

      @os-trump

      Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

      What

      Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

      Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

      1192 uncaughtException code=EPIPE msg=write EPIPE
      stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
      | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
      1192 exit code=1
      

      12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

      Why it matters

      Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

      parent's readerexitelapsed
      drained2~8.3 s (145696 bytes delivered)
      paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
      read end destroyed11.4 s

      So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

      The second consequence: writeStderr() is never reached

      bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

      Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

      bin/run-dev.js stateexitelapsed
      pristine11387-1711 ms
      write callback removed, so the closed path can only finish on the bound11517-1633 ms
      process.stderr.on('error', noop) added, callback kept28787-8979 ms
      both223601-23712 ms (one bound paid)

      Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

      What a fix owes

      1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
      2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
      3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

      Related

      Generated by Claude Code

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

          Description

          @os-trump

          Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

          What

          Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

          Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

          1192 uncaughtException code=EPIPE msg=write EPIPE
          stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
          | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
          1192 exit code=1
          

          12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

          Why it matters

          Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

          parent's readerexitelapsed
          drained2~8.3 s (145696 bytes delivered)
          paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
          read end destroyed11.4 s

          So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

          The second consequence: writeStderr() is never reached

          bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

          Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

          bin/run-dev.js stateexitelapsed
          pristine11387-1711 ms
          write callback removed, so the closed path can only finish on the bound11517-1633 ms
          process.stderr.on('error', noop) added, callback kept28787-8979 ms
          both223601-23712 ms (one bound paid)

          Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

          What a fix owes

          1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
          2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
          3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

          Related

          Generated by Claude Code

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

              Description

              @os-trump

              Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

              What

              Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

              Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

              1192 uncaughtException code=EPIPE msg=write EPIPE
              stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
              | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
              1192 exit code=1
              

              12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

              Why it matters

              Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

              parent's readerexitelapsed
              drained2~8.3 s (145696 bytes delivered)
              paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
              read end destroyed11.4 s

              So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

              The second consequence: writeStderr() is never reached

              bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

              Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

              bin/run-dev.js stateexitelapsed
              pristine11387-1711 ms
              write callback removed, so the closed path can only finish on the bound11517-1633 ms
              process.stderr.on('error', noop) added, callback kept28787-8979 ms
              both223601-23712 ms (one bound paid)

              Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

              What a fix owes

              1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
              2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
              3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

              Related

              Generated by Claude Code

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

                  Description

                  @os-trump

                  Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

                  What

                  Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

                  Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

                  1192 uncaughtException code=EPIPE msg=write EPIPE
                  stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
                  | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
                  1192 exit code=1
                  

                  12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

                  Why it matters

                  Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

                  parent's readerexitelapsed
                  drained2~8.3 s (145696 bytes delivered)
                  paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
                  read end destroyed11.4 s

                  So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

                  The second consequence: writeStderr() is never reached

                  bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

                  Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

                  bin/run-dev.js stateexitelapsed
                  pristine11387-1711 ms
                  write callback removed, so the closed path can only finish on the bound11517-1633 ms
                  process.stderr.on('error', noop) added, callback kept28787-8979 ms
                  both223601-23712 ms (one bound paid)

                  Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

                  What a fix owes

                  1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
                  2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
                  3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

                  Related

                  Generated by Claude Code

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

                      Description

                      @os-trump

                      Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

                      What

                      Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

                      Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

                      1192 uncaughtException code=EPIPE msg=write EPIPE
                      stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
                      | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
                      1192 exit code=1
                      

                      12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

                      Why it matters

                      Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

                      parent's readerexitelapsed
                      drained2~8.3 s (145696 bytes delivered)
                      paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
                      read end destroyed11.4 s

                      So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

                      The second consequence: writeStderr() is never reached

                      bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

                      Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

                      bin/run-dev.js stateexitelapsed
                      pristine11387-1711 ms
                      write callback removed, so the closed path can only finish on the bound11517-1633 ms
                      process.stderr.on('error', noop) added, callback kept28787-8979 ms
                      both223601-23712 ms (one bound paid)

                      Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

                      What a fix owes

                      1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
                      2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
                      3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

                      Related

                      Generated by Claude Code

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

                          Description

                          @os-trump

                          Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

                          What

                          Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

                          Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

                          1192 uncaughtException code=EPIPE msg=write EPIPE
                          stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
                          | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
                          1192 exit code=1
                          

                          12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

                          Why it matters

                          Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

                          parent's readerexitelapsed
                          drained2~8.3 s (145696 bytes delivered)
                          paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
                          read end destroyed11.4 s

                          So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

                          The second consequence: writeStderr() is never reached

                          bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

                          Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

                          bin/run-dev.js stateexitelapsed
                          pristine11387-1711 ms
                          write callback removed, so the closed path can only finish on the bound11517-1633 ms
                          process.stderr.on('error', noop) added, callback kept28787-8979 ms
                          both223601-23712 ms (one bound paid)

                          Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

                          What a fix owes

                          1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
                          2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
                          3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

                          Related

                          Generated by Claude Code

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              os dev dies of an uncaught write EPIPE (exit 1) when its stderr read end is CLOSED — every other reader gets exit 2, and the drain is never reached #14858

                              Description

                              @os-trump

                              Found while fixing #14716 (the closed-read-end oracle in packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts). Not fixed there: that card's file fence is the test file, and the fix for this is in packages/cli/bin/run-dev.js. Filed unassigned and unlabelled for triage.

                              What

                              Spawn the CLI with its stderr piped and destroy the parent's read end (stdio: ['ignore','ignore','pipe'], then child.stderr.destroy()). oclif's displayWarnings() makes the first stderr write, the pipe is already gone, node raises write EPIPE as an error event on process.stderr, nothing is listening, and the process dies of an uncaught exception: exit 1, ~1.4 s in.

                              Traced with a --import observer that only appends to a file (no listener on process.stderr, no write wrapper), 3 for 3:

                              1192 uncaughtException code=EPIPE msg=write EPIPE
                              stack=Error: write EPIPE | at afterWriteDispatched (node:internal/stream_base_commons:159:15)
                              | at writeGeneric (...) | at Socket._writeGeneric (node:net:966:11)
                              1192 exit code=1
                              

                              12 further untraced iterations on the same box: 12/12 code=1 signal=null, 1387-1711 ms.

                              Why it matters

                              Every other reader of the same child gets exit 2 — the status oclif uses for a failed command, and the status #14715 pinned for the never-read reader. Measured here with the same probe and flags:

                              parent's readerexitelapsed
                              drained2~8.3 s (145696 bytes delivered)
                              paused, never read28.4 s / 23.5 s (the second paid the shim's 15 s bound)
                              read end destroyed11.4 s

                              So a caller that closes stderr — a supervisor that drops the pipe, os dev 2 with the fd closed — cannot tell "the command failed" from "the CLI crashed", and the exit status is the only channel it has left.

                              The second consequence: writeStderr() is never reached

                              bin/run-dev.js has an elaborate drain with a no-progress bound, and its comment says the write callback "also fires on EPIPE, which is what releases the closed-reader paths promptly." That is true of node, but the closed-reader path never gets there — the process is dead ~7.8 s earlier. Traced with an observer that DOES swallow the error event, the shim's own 415-byte write is write #175, at 9250 ms, behind 174 oclif writes that all EPIPE.

                              Confirmed by ablation on a scratch worktree (mutations verified on disk by blob hash, restored by blob-hash equality under a trap):

                              bin/run-dev.js stateexitelapsed
                              pristine11387-1711 ms
                              write callback removed, so the closed path can only finish on the bound11517-1633 ms
                              process.stderr.on('error', noop) added, callback kept28787-8979 ms
                              both223601-23712 ms (one bound paid)

                              Line 3 is the whole finding: one noop error listener is the difference between exit 1 at 1.4 s and the contracted exit 2 at 8.8 s, and only in that state does the shim's carefully written drain run at all for a closed reader.

                              What a fix owes

                              1. Decide whether a closed stderr should still exit 2. If yes, the CLI needs an EPIPE-tolerant process.stderr (a noop error listener is the smallest form, but it changes what the drain does for this path, so it wants the ablation table above re-measured rather than assumed).
                              2. ⚠️ Whatever lands, packages/cli/test/run-dev-unbuilt-workspace.e2e.test.ts pins the CURRENT status (closedEnd.code is 1) precisely so this cannot change silently. The fixing PR flips that number and says why — it is not a broken test.
                              3. ⛔ Not the same defect as os dev in an unbuilt workspace sometimes HANGS instead of exiting 2 when its reader goes away — measured at 180103 ms against a 7046 ms calibration on the same runner #14832. That one is an intermittent HANG of the never-read (paused) reader, 180103 ms against a 7046 ms calibration; this one is deterministic, 12/12, and specific to a DESTROYED read end. Same file, opposite failure modes. Do not merge them.

                              Related

                              Generated by Claude Code

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions