autobuild: relaunch after install is a race, and a failed relaunch is silent #41

Description

@radroid

Symptom

On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
opened by hand. Every status indicator stayed green throughout.

Evidence chain

TimeEventSource
01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
01:03:44last trace record written by the old appdesktop.trace.ndjson
01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

Root cause

:401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
handler only reveals the existing window), so open activated the dying instance and
returned 0.

The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
now a coin flip; 07-30 won it, 08-02 lost it.

Two further causes of the same symptom, independent of the race

  • The bundle is rm -rf'd under six live processes (main, server child,
    t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
    app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
    path that is missing or is the new build under the old Electron main.
  • Port squat → silent green. If that child outlives the quit, the new app walks up to
    :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
  • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
    launch, open still exits 0.

Fixes, ranked

All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

  1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
    wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
    returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
    timeout abandon the staged copy and return 1without touching the installed app. Only
    then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
    child respawn and the port squat in one change. Reject open -n — the single-instance
    lock makes a second instance quit immediately after revealing the dying one.
  2. Verify the relaunch and make it reportable. Require a new pid; poll
    curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
    contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
    appPid, backendPort and a distinct installed-not-running result to the status JSON.
    Nothing consumes that file today, so this is purely additive.
  3. Post-install launch watchdog — see spec below.

Watchdog spec (as decided)

A separate script, spawned by the autobuild once the install completes, detached so it
survives the parent:

  • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
  • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
    never linger.
  • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
  • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
    pgrep — otherwise a port squat reads as healthy.

Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
with 10 minutes a one-line change.

Known gap: post-install-only means it does not catch the app dying for any other reason
(crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
the failure actually observed. No deferral for running Claude sessions — installs proceed and
the app is reopened afterwards.

Two traps to fix at the same time

  • --print-launchd would silently undo the cargo fix. The generator hardcodes
    PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
    grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
    and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
    back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
    resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
  • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
    interval default as 60, and :200 — the --print-launchd example itself — uses
    --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
    agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
    sessions.

Smaller items

KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
written into the tree the sync branch lives in); atomic write_status; -readonly on the
install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
trap; reject --relaunch without --install; self-heal when the app is missing (the marker
is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
build" forever).

Metadata

Metadata

Assignees

No one assigned

    Labels

    wontfixThis will not be worked on

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      autobuild: relaunch after install is a race, and a failed relaunch is silent #41

      Description

      @radroid

      Symptom

      On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
      The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
      opened by hand. Every status indicator stayed green throughout.

      Evidence chain

      TimeEventSource
      01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
      01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
      01:03:44last trace record written by the old appdesktop.trace.ndjson
      01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
      01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
      02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

      Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

      Root cause

      :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
      cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
      still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
      handler only reveals the existing window), so open activated the dying instance and
      returned 0.

      The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
      12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
      now a coin flip; 07-30 won it, 08-02 lost it.

      Two further causes of the same symptom, independent of the race

      • The bundle is rm -rf'd under six live processes (main, server child,
        t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
        app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
        path that is missing or is the new build under the old Electron main.
      • Port squat → silent green. If that child outlives the quit, the new app walks up to
        :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
      • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
        launch, open still exits 0.

      Fixes, ranked

      All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

      1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
        wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
        returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
        timeout abandon the staged copy and return 1without touching the installed app. Only
        then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
        child respawn and the port squat in one change. Reject open -n — the single-instance
        lock makes a second instance quit immediately after revealing the dying one.
      2. Verify the relaunch and make it reportable. Require a new pid; poll
        curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
        contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
        appPid, backendPort and a distinct installed-not-running result to the status JSON.
        Nothing consumes that file today, so this is purely additive.
      3. Post-install launch watchdog — see spec below.

      Watchdog spec (as decided)

      A separate script, spawned by the autobuild once the install completes, detached so it
      survives the parent:

      • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
      • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
        never linger.
      • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
      • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
        pgrep — otherwise a port squat reads as healthy.

      Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
      10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
      with 10 minutes a one-line change.

      Known gap: post-install-only means it does not catch the app dying for any other reason
      (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
      the failure actually observed. No deferral for running Claude sessions — installs proceed and
      the app is reopened afterwards.

      Two traps to fix at the same time

      • --print-launchd would silently undo the cargo fix. The generator hardcodes
        PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
        grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
        and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
        back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
        resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
      • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
        interval default as 60, and :200 — the --print-launchd example itself — uses
        --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
        agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
        sessions.

      Smaller items

      KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
      indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
      install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
      written into the tree the sync branch lives in); atomic write_status; -readonly on the
      install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
      trap; reject --relaunch without --install; self-heal when the app is missing (the marker
      is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
      build" forever).

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        wontfixThis will not be worked on

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          autobuild: relaunch after install is a race, and a failed relaunch is silent #41

          Description

          @radroid

          Symptom

          On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
          The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
          opened by hand. Every status indicator stayed green throughout.

          Evidence chain

          TimeEventSource
          01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
          01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
          01:03:44last trace record written by the old appdesktop.trace.ndjson
          01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
          01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
          02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

          Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

          Root cause

          :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
          cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
          still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
          handler only reveals the existing window), so open activated the dying instance and
          returned 0.

          The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
          12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
          now a coin flip; 07-30 won it, 08-02 lost it.

          Two further causes of the same symptom, independent of the race

          • The bundle is rm -rf'd under six live processes (main, server child,
            t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
            app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
            path that is missing or is the new build under the old Electron main.
          • Port squat → silent green. If that child outlives the quit, the new app walks up to
            :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
          • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
            launch, open still exits 0.

          Fixes, ranked

          All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

          1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
            wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
            returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
            timeout abandon the staged copy and return 1without touching the installed app. Only
            then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
            child respawn and the port squat in one change. Reject open -n — the single-instance
            lock makes a second instance quit immediately after revealing the dying one.
          2. Verify the relaunch and make it reportable. Require a new pid; poll
            curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
            contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
            appPid, backendPort and a distinct installed-not-running result to the status JSON.
            Nothing consumes that file today, so this is purely additive.
          3. Post-install launch watchdog — see spec below.

          Watchdog spec (as decided)

          A separate script, spawned by the autobuild once the install completes, detached so it
          survives the parent:

          • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
          • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
            never linger.
          • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
          • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
            pgrep — otherwise a port squat reads as healthy.

          Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
          10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
          with 10 minutes a one-line change.

          Known gap: post-install-only means it does not catch the app dying for any other reason
          (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
          the failure actually observed. No deferral for running Claude sessions — installs proceed and
          the app is reopened afterwards.

          Two traps to fix at the same time

          • --print-launchd would silently undo the cargo fix. The generator hardcodes
            PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
            grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
            and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
            back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
            resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
          • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
            interval default as 60, and :200 — the --print-launchd example itself — uses
            --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
            agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
            sessions.

          Smaller items

          KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
          indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
          install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
          written into the tree the sync branch lives in); atomic write_status; -readonly on the
          install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
          trap; reject --relaunch without --install; self-heal when the app is missing (the marker
          is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
          build" forever).

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            wontfixThis will not be worked on

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              autobuild: relaunch after install is a race, and a failed relaunch is silent #41

              Description

              @radroid

              Symptom

              On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
              The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
              opened by hand. Every status indicator stayed green throughout.

              Evidence chain

              TimeEventSource
              01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
              01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
              01:03:44last trace record written by the old appdesktop.trace.ndjson
              01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
              01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
              02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

              Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

              Root cause

              :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
              cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
              still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
              handler only reveals the existing window), so open activated the dying instance and
              returned 0.

              The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
              12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
              now a coin flip; 07-30 won it, 08-02 lost it.

              Two further causes of the same symptom, independent of the race

              • The bundle is rm -rf'd under six live processes (main, server child,
                t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
                app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
                path that is missing or is the new build under the old Electron main.
              • Port squat → silent green. If that child outlives the quit, the new app walks up to
                :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
              • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
                launch, open still exits 0.

              Fixes, ranked

              All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

              1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
                wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
                returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
                timeout abandon the staged copy and return 1without touching the installed app. Only
                then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
                child respawn and the port squat in one change. Reject open -n — the single-instance
                lock makes a second instance quit immediately after revealing the dying one.
              2. Verify the relaunch and make it reportable. Require a new pid; poll
                curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
                contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
                appPid, backendPort and a distinct installed-not-running result to the status JSON.
                Nothing consumes that file today, so this is purely additive.
              3. Post-install launch watchdog — see spec below.

              Watchdog spec (as decided)

              A separate script, spawned by the autobuild once the install completes, detached so it
              survives the parent:

              • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
              • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
                never linger.
              • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
              • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
                pgrep — otherwise a port squat reads as healthy.

              Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
              10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
              with 10 minutes a one-line change.

              Known gap: post-install-only means it does not catch the app dying for any other reason
              (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
              the failure actually observed. No deferral for running Claude sessions — installs proceed and
              the app is reopened afterwards.

              Two traps to fix at the same time

              • --print-launchd would silently undo the cargo fix. The generator hardcodes
                PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
                grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
                and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
                back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
                resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
              • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
                interval default as 60, and :200 — the --print-launchd example itself — uses
                --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
                agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
                sessions.

              Smaller items

              KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
              indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
              install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
              written into the tree the sync branch lives in); atomic write_status; -readonly on the
              install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
              trap; reject --relaunch without --install; self-heal when the app is missing (the marker
              is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
              build" forever).

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                wontfixThis will not be worked on

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  autobuild: relaunch after install is a race, and a failed relaunch is silent #41

                  Description

                  @radroid

                  Symptom

                  On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
                  The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
                  opened by hand. Every status indicator stayed green throughout.

                  Evidence chain

                  TimeEventSource
                  01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
                  01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
                  01:03:44last trace record written by the old appdesktop.trace.ndjson
                  01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
                  01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
                  02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

                  Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

                  Root cause

                  :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
                  cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
                  still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
                  handler only reveals the existing window), so open activated the dying instance and
                  returned 0.

                  The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
                  12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
                  now a coin flip; 07-30 won it, 08-02 lost it.

                  Two further causes of the same symptom, independent of the race

                  • The bundle is rm -rf'd under six live processes (main, server child,
                    t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
                    app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
                    path that is missing or is the new build under the old Electron main.
                  • Port squat → silent green. If that child outlives the quit, the new app walks up to
                    :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
                  • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
                    launch, open still exits 0.

                  Fixes, ranked

                  All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

                  1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
                    wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
                    returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
                    timeout abandon the staged copy and return 1without touching the installed app. Only
                    then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
                    child respawn and the port squat in one change. Reject open -n — the single-instance
                    lock makes a second instance quit immediately after revealing the dying one.
                  2. Verify the relaunch and make it reportable. Require a new pid; poll
                    curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
                    contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
                    appPid, backendPort and a distinct installed-not-running result to the status JSON.
                    Nothing consumes that file today, so this is purely additive.
                  3. Post-install launch watchdog — see spec below.

                  Watchdog spec (as decided)

                  A separate script, spawned by the autobuild once the install completes, detached so it
                  survives the parent:

                  • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
                  • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
                    never linger.
                  • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
                  • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
                    pgrep — otherwise a port squat reads as healthy.

                  Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
                  10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
                  with 10 minutes a one-line change.

                  Known gap: post-install-only means it does not catch the app dying for any other reason
                  (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
                  the failure actually observed. No deferral for running Claude sessions — installs proceed and
                  the app is reopened afterwards.

                  Two traps to fix at the same time

                  • --print-launchd would silently undo the cargo fix. The generator hardcodes
                    PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
                    grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
                    and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
                    back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
                    resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
                  • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
                    interval default as 60, and :200 — the --print-launchd example itself — uses
                    --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
                    agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
                    sessions.

                  Smaller items

                  KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
                  indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
                  install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
                  written into the tree the sync branch lives in); atomic write_status; -readonly on the
                  install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
                  trap; reject --relaunch without --install; self-heal when the app is missing (the marker
                  is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
                  build" forever).

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    wontfixThis will not be worked on

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      autobuild: relaunch after install is a race, and a failed relaunch is silent #41

                      Description

                      @radroid

                      Symptom

                      On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
                      The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
                      opened by hand. Every status indicator stayed green throughout.

                      Evidence chain

                      TimeEventSource
                      01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
                      01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
                      01:03:44last trace record written by the old appdesktop.trace.ndjson
                      01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
                      01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
                      02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

                      Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

                      Root cause

                      :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
                      cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
                      still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
                      handler only reveals the existing window), so open activated the dying instance and
                      returned 0.

                      The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
                      12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
                      now a coin flip; 07-30 won it, 08-02 lost it.

                      Two further causes of the same symptom, independent of the race

                      • The bundle is rm -rf'd under six live processes (main, server child,
                        t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
                        app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
                        path that is missing or is the new build under the old Electron main.
                      • Port squat → silent green. If that child outlives the quit, the new app walks up to
                        :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
                      • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
                        launch, open still exits 0.

                      Fixes, ranked

                      All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

                      1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
                        wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
                        returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
                        timeout abandon the staged copy and return 1without touching the installed app. Only
                        then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
                        child respawn and the port squat in one change. Reject open -n — the single-instance
                        lock makes a second instance quit immediately after revealing the dying one.
                      2. Verify the relaunch and make it reportable. Require a new pid; poll
                        curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
                        contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
                        appPid, backendPort and a distinct installed-not-running result to the status JSON.
                        Nothing consumes that file today, so this is purely additive.
                      3. Post-install launch watchdog — see spec below.

                      Watchdog spec (as decided)

                      A separate script, spawned by the autobuild once the install completes, detached so it
                      survives the parent:

                      • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
                      • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
                        never linger.
                      • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
                      • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
                        pgrep — otherwise a port squat reads as healthy.

                      Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
                      10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
                      with 10 minutes a one-line change.

                      Known gap: post-install-only means it does not catch the app dying for any other reason
                      (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
                      the failure actually observed. No deferral for running Claude sessions — installs proceed and
                      the app is reopened afterwards.

                      Two traps to fix at the same time

                      • --print-launchd would silently undo the cargo fix. The generator hardcodes
                        PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
                        grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
                        and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
                        back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
                        resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
                      • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
                        interval default as 60, and :200 — the --print-launchd example itself — uses
                        --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
                        agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
                        sessions.

                      Smaller items

                      KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
                      indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
                      install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
                      written into the tree the sync branch lives in); atomic write_status; -readonly on the
                      install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
                      trap; reject --relaunch without --install; self-heal when the app is missing (the marker
                      is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
                      build" forever).

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        wontfixThis will not be worked on

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          autobuild: relaunch after install is a race, and a failed relaunch is silent #41

                          Description

                          @radroid

                          Symptom

                          On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
                          The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
                          opened by hand. Every status indicator stayed green throughout.

                          Evidence chain

                          TimeEventSource
                          01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
                          01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
                          01:03:44last trace record written by the old appdesktop.trace.ndjson
                          01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
                          01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
                          02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

                          Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

                          Root cause

                          :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
                          cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
                          still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
                          handler only reveals the existing window), so open activated the dying instance and
                          returned 0.

                          The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
                          12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
                          now a coin flip; 07-30 won it, 08-02 lost it.

                          Two further causes of the same symptom, independent of the race

                          • The bundle is rm -rf'd under six live processes (main, server child,
                            t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
                            app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
                            path that is missing or is the new build under the old Electron main.
                          • Port squat → silent green. If that child outlives the quit, the new app walks up to
                            :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
                          • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
                            launch, open still exits 0.

                          Fixes, ranked

                          All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

                          1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
                            wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
                            returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
                            timeout abandon the staged copy and return 1without touching the installed app. Only
                            then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
                            child respawn and the port squat in one change. Reject open -n — the single-instance
                            lock makes a second instance quit immediately after revealing the dying one.
                          2. Verify the relaunch and make it reportable. Require a new pid; poll
                            curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
                            contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
                            appPid, backendPort and a distinct installed-not-running result to the status JSON.
                            Nothing consumes that file today, so this is purely additive.
                          3. Post-install launch watchdog — see spec below.

                          Watchdog spec (as decided)

                          A separate script, spawned by the autobuild once the install completes, detached so it
                          survives the parent:

                          • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
                          • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
                            never linger.
                          • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
                          • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
                            pgrep — otherwise a port squat reads as healthy.

                          Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
                          10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
                          with 10 minutes a one-line change.

                          Known gap: post-install-only means it does not catch the app dying for any other reason
                          (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
                          the failure actually observed. No deferral for running Claude sessions — installs proceed and
                          the app is reopened afterwards.

                          Two traps to fix at the same time

                          • --print-launchd would silently undo the cargo fix. The generator hardcodes
                            PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
                            grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
                            and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
                            back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
                            resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
                          • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
                            interval default as 60, and :200 — the --print-launchd example itself — uses
                            --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
                            agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
                            sessions.

                          Smaller items

                          KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
                          indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
                          install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
                          written into the tree the sync branch lives in); atomic write_status; -readonly on the
                          install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
                          trap; reject --relaunch without --install; self-heal when the app is missing (the marker
                          is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
                          build" forever).

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            wontfixThis will not be worked on

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              autobuild: relaunch after install is a race, and a failed relaunch is silent #41

                              Description

                              @radroid

                              Symptom

                              On 2026-08-02 the autobuild built and installed aaedc21ab, then left nothing running.
                              The app and :3773 (the Tailscale remote surface) were down for 103 minutes, until it was
                              opened by hand. Every status indicator stayed green throughout.

                              Evidence chain

                              TimeEventSource
                              01:03:41osascript quit sent; bundle replaced (mtime 01:03:41–42)auto-build-desktop.sh:401, stat
                              01:03:42log: install: relaunching; open runs; lsd re-registers the bundle — no process launchedunified log
                              01:03:44last trace record written by the old appdesktop.trace.ndjson
                              01:03:45.165launchd: application.com.t3tools.t3code…[73031] exited due to exit(0), ran for 216120549msunified log
                              01:03:45 → 02:46:40zero t3code launch activity, zero trace recordsunified log + trace
                              02:46:40launch originated by com.apple.coreservices.uiagent — opened by handunified log

                              Status file said {"result":"built","installed":true,"detail":"ok"} the entire time.

                              Root cause

                              :401 fires an asynchronous AppleScript quit; :431 calls open. Between them sit only
                              cp -R, rm -rf, mv, xattr — nothing waits for the old process to exit. The old instance
                              still held Electron's single-instance lock (DesktopClerk.ts:90,128, whose second-instance
                              handler only reveals the existing window), so open activated the dying instance and
                              returned 0.

                              The cp -R was the accidental sleep that made this work at all. Copy windows: 10s (07-27),
                              12s, 12s, then 4s (07-30) and 1s (08-02) as the artifact shrank 237MB → 150MB. The race is
                              now a coin flip; 07-30 won it, 08-02 lost it.

                              Two further causes of the same symptom, independent of the race

                              • The bundle is rm -rf'd under six live processes (main, server child,
                                t3-resource-monitor, three helpers). :3773 is owned by the server child, not the main
                                app, and DesktopBackendManager has a restart loop — a child killed mid-swap respawns from a
                                path that is missing or is the new build under the old Electron main.
                              • Port squat → silent green. If that child outlives the quit, the new app walks up to
                                :3774 while tailscale serve is pinned to 3773. Identical outage, app fully alive.
                              • xattr -dr com.apple.quarantine … || true (:429) swallows failure — Gatekeeper refuses the
                                launch, open still exits 0.

                              Fixes, ranked

                              All in scripts/t3x/** and docs/t3x/** — fork-owned, zero upstream churn, no rebase cost.

                              1. Reorder install_dmg.cp -R first; capture the quit's output instead of discarding it;
                                wait_for_exit polling ps -Ao pid=,comm= (pid=,command= matches its own pipeline and
                                returns 9 pids instead of 6 — it would time out on every install); re-quit → SIGTERM; on
                                timeout abandon the staged copy and return 1without touching the installed app. Only
                                then rm/mv/xattr/open. Fixes the race, the replace-under-live-process hazard, the
                                child respawn and the port squat in one change. Reject open -n — the single-instance
                                lock makes a second instance quit immediately after revealing the dying one.
                              2. Verify the relaunch and make it reportable. Require a new pid; poll
                                curl -fsS http://127.0.0.1:3773/.well-known/t3/environment (the app's own readiness
                                contract); assert the :3773 listener is in the bundle's pid set. Add relaunched,
                                appPid, backendPort and a distinct installed-not-running result to the status JSON.
                                Nothing consumes that file today, so this is purely additive.
                              3. Post-install launch watchdog — see spec below.

                              Watchdog spec (as decided)

                              A separate script, spawned by the autobuild once the install completes, detached so it
                              survives the parent:

                              • Wait a grace period, then poll every 5 minutes: if the app is not running, open it.
                              • Exit as soon as the app is confirmed up. Exit after a bounded number of attempts so it can
                                never linger.
                              • Skip while the autobuild lock is held, so it cannot fight a concurrent install.
                              • Readiness = the :3773 probe and the listener pid belonging to the bundle, not just
                                pgrep — otherwise a port squat reads as healthy.

                              Grace period: requested as 10 minutes. The app starts in ~5 seconds, so 10 minutes is up to
                              10 minutes of dead :3773 in the common case. Suggest a named constant defaulting to ~90s,
                              with 10 minutes a one-line change.

                              Known gap: post-install-only means it does not catch the app dying for any other reason
                              (crash, manual quit, OOM). A standing LaunchAgent would; this is the simpler shape and covers
                              the failure actually observed. No deferral for running Claude sessions — installs proceed and
                              the app is reopened afterwards.

                              Two traps to fix at the same time

                              • --print-launchd would silently undo the cargo fix. The generator hardcodes
                                PATH=/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin (:625). grep -c cargo and
                                grep -c KeepAlive on the script both return 0. The live plist carries the fnm node path
                                and~/.cargo/bin — the 07-30 spawn cargo ENOENT fix was hand-applied and never
                                back-ported. Regenerating today reinstates a 12-hour silent build outage. Derive PATH from
                                resolved node/pnpm/cargo/git, refuse to emit if any is missing, add --diff-launchd.
                              • The runbook will install a 2-minute loop.auto-build-runbook.md:127 documents the
                                interval default as 60, and :200 — the --print-launchd example itself — uses
                                --interval 120. The script is INTERVAL=43200 (:76). Following the docs to rebuild the
                                agent produces a two-minute quit/replace/relaunch cycle on an app that hosts live Claude
                                sessions.

                              Smaller items

                              KeepAlive + ThrottleInterval on the agent (if the watcher dies nothing restarts it and every
                              indicator stays green); dmg freshness check (newest_dmg is an mtime pick, so a stale build can
                              install fully green); move T3CODE_DESKTOP_OUTPUT_DIR off the main checkout (2.4GB currently
                              written into the tree the sync branch lives in); atomic write_status; -readonly on the
                              install-path hdiutil attach (:391 lacks it, the peek path :338 has it); INT TERM detach
                              trap; reject --relaunch without --install; self-heal when the app is missing (the marker
                              is a built-sha, not an installed-state, so deleting the bundle yields "no change since last
                              build" forever).

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                wontfixThis will not be worked on

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions