perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

perf(ci): cut about a minute from every release - #110

Merged
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf
Aug 27, 2026
Merged

perf(ci): cut about a minute from every release#110
rynfar merged 2 commits into
pylonfrom
upstream/2026-08-27-release-ci-perf

Conversation

@rynfar

@rynfarrynfar commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

Two pieces of every release run are pure waste: publish_cli builds the web client twice, and two startup jobs wait on version resolution they never read. This removes both, and fixes the two scheduling problems upstream found in the same audit.

Pylon's release workflow has diverged too far from upstream's to cherry-pick, so this is a manual port of pingdotgg/t3code#8250 (a3a8cbd60).

What changed

  • publish_cli built the web client twice.@t3tools/web is a workspace dependency of t3, so vp run --filter t3 build already builds it. The explicit step repeated the work on the serial publish path. Verified locally: the command runs exactly two tasks (web, then server) and still produces dist/bin.mjs, dist/service-launcher.mjs, and dist/client/index.html.
  • relay_public_config and build_wsl_node_pty now start alongside preflight. Both consume only the commit SHA, and Pylon's preflight.outputs.ref is literally ${{ github.sha }}, so the swap is exactly equivalent. They now share preflight's own check_changes condition. Desktop builds still hard-gate on needs.preflight.result == 'success' && needs.relay_public_config.result == 'success', so nothing builds off a failed preflight.
  • Nightly cron moves to minute 7. GitHub delays scheduled runs most at the top of the hour; upstream measured a 13-minute median start delay across 56 sampled runs.
  • Nightlies queue instead of cancelling. See below.

Why the concurrency policy flips

8dd0d6e06 set cancel-in-progress: ${{ github.event_name == 'schedule' }} after a Blacksmith capacity outage let 23 nightlies stack ~480 VM-hours deep waiting for runners that never arrived. That was the right fix at the time. 124630c3f then moved every workflow onto GitHub-hosted runners yesterday, and the concurrency block was never revisited — its comment still describes a failure mode that can no longer occur.

Cancelling is now the more dangerous half. publish_cli pushes to npm beforerelease creates the GitHub release, so a cancel in that window strands a published CLI version with no matching desktop build, and npm will not take that version number again. cancel-in-progress: false closes it, and queue: max keeps every pending run (100 FIFO slots per group) instead of the default newest-wins single slot, so the "losing a real release to a scheduling race" concern that motivated the run_id grouping is still covered — a queued stable tag waits rather than being dropped. Stable releases now serialize rather than run in parallel, which at Pylon's cadence is the safer trade.

The two settings are not separable: queue: max with cancel-in-progress: true is a workflow validation error.

Checks

  • actionlint 1.7.12: 4 findings on pylon today, 5 on this branch. The one new finding is unexpected key "queue" for "concurrency" section — the linter's schema predates the feature, which GitHub shipped in May 2026. No Pylon workflow runs actionlint in CI, so it does not gate anything.
  • The other 4 findings (run_started_at, three shellcheck SC2129) are unchanged from pylon.

One behavior change worth noting

relay_public_config no longer skips when preflight fails, so a failed preflight now costs one extra ~5-minute job that resolves production relay config and then goes unused. Upstream accepted the same trade.

Model: Claude Opus 5. Harness: Claude Code.


View with [code]smithAutofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Two pieces of every release run were pure waste: publish_cli built the web
client twice, and two startup jobs waited on version resolution they never
read. Fixes both, plus the two scheduling problems upstream found in the
same audit.
Adapted from pingdotgg/t3code#8250 (a3a8cbd60). Pylon's release workflow has
diverged too far to cherry-pick, so this is a manual port. The concurrency
comment is rewritten rather than transplanted so Pylon's own incident history
stays recorded.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 27, 2026
@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.3 KiB13.3 KiB+60 B (+0.4%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB+9 B (+0.1%)7.3 KiB
CodexLive turn WebSocket wire6.4 KiB6.4 KiB+51 B (+0.8%)7.8 KiB
CodexLive turn WebSocket decoded55.5 KiB55.6 KiB+44 B (+0.1%)66.4 KiB
CodexLive turn messages910+1 (+11.1%)21
ClaudeTotal thread wire13.3 KiB13.3 KiB+1 B (+0.0%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−8 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.4 KiB+9 B (+0.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.4 KiB0 B (0.0%)66.4 KiB
ClaudeLive turn messages10100 (0.0%)21

Baseline: 3ba1857 · PR result: 8d17396 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.0 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Adversarial review of the port found two problems with taking upstream's
concurrency block as written.
Upstream collapses every tag push and manual stable dispatch into one shared
`release-stable` group. Upstream had no concurrency block at all before this,
so that was strictly more serialization from zero; Pylon has keyed stable runs
off run_id since 8dd0d6e specifically so they never interact. With
cancel-in-progress now false, adopting the shared lane would queue a P0 hotfix
tag behind an in-flight release for the length of its build matrix. Stable
keeps run_id; nightlies keep the shared group they need.
check_changes was also the only job in the file without timeout-minutes. That
was survivable while a fresh nightly cancelled a wedged predecessor. Now that
nightlies queue, a stuck checkout there would hold release-nightly across two
cron slots until GitHub's 360-minute default fired.
@rynfar

Copy link
Copy Markdown
CollaboratorAuthor

Adversarial review before merge turned up two real problems with taking upstream's concurrency block verbatim. Both fixed in 8d1739684.

Stable releases were losing per-run isolation. Upstream's group collapses every tag push and manual stable dispatch into one literal release-stable lane. Upstream had no concurrency block before #8250, so for them that was strictly more serialization from zero — but Pylon has keyed stable runs off run_id since 8dd0d6e06 precisely so they never interact. Combined with cancel-in-progress: false, adopting the shared lane means a P0 hotfix tag pushed while a release is in flight goes pending behind that run's 90-minute build matrix and 30-minute release upload. The group now keeps github.run_id for stable and the shared group only for nightlies, which preserves both properties:

group: release-${{ (github.event_name == 'schedule' || inputs.channel == 'nightly') && 'nightly' || github.run_id }}

check_changes had no timeout-minutes. It was the only job in the file relying on GitHub's 360-minute default. That was survivable while a fresh nightly cancelled a wedged predecessor; with queuing, a stuck checkout there would hold release-nightly across two cron slots. Now bounded at 5 minutes.

One finding left as-is, deliberately: relay_public_config and build_wsl_node_pty no longer skip when preflight fails, so a doomed run still spends up to 30 minutes on node-gyp rebuild. That is the cost of the parallelism and upstream accepted it. Its blast radius drops a lot with the two fixes above — stable runs no longer hold a shared lane at all, and the nightly case is bounded well inside the 3-hour cron gap. Verified production has protection_rules: [] and no deployment branch policy, so there is no approval gate or secret exposure, just wasted minutes.

Also confirmed clean under review: inputs.channel is legal in the concurrency context and resolves to the right group for every event type; preflight.outputs.ref really is literally ${{ github.sha }}, so that swap is a textual no-op; and apps/server/vite.config.ts declares dependsOn: ["@t3tools/web#build"], so deleting the explicit web build genuinely removes duplicated work rather than dropping it.

@rynfar
rynfar merged commit 6ca228f into pylonAug 27, 2026
15 checks passed
@rynfar
rynfar deleted the upstream/2026-08-27-release-ci-perf branch August 27, 2026 06:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@rynfar