test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1) - #522

Merged
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift
Jun 5, 2026
Merged

test(e2e): pin live weight-edit re-takes effect on dispatch (#196 L1)#522
moonming merged 1 commit into
mainfrom
test/issue-L1a-weight-edit-shift

Conversation

@moonming

Copy link
Copy Markdown
Member

What

Closes the L1 gap (config edit → observable behavior change). The sibling weighted-routing-distribution-e2e pins that a weighted model's initial weights are honored. This pins that a live edit to those weights propagates and the weighted scheduler rebuilds — a scheduler that cached its weight-wheel on first dispatch and ignored config updates would silently keep serving the old split.

How (deterministic, no statistics)

weight 0 = excluded (per routing-strategies-e2e "weighted picks the positive-weight target"):

  1. Create wr-edit-virtual weighted [wr-edit-a: 100, wr-edit-b: 0] → assert all BATCH dispatches hit A.
  2. PUT /admin/v1/models/:id inverting to [0, 100].
  3. Propagation signal: a virtual probe returning "served by B"impossible under the old [100,0] config, so it proves the edit is live and the scheduler rebuilt before counting.
  4. Assert all BATCH dispatches now hit B.

If the scheduler never rebuilds on a config edit (the regression this targets), step 3's waitConfigPropagation times out — surfacing it loudly rather than passing silently.

Note on a possible real finding

If CI shows step 3 timing out, that's a real DP bug (weighted scheduler doesn't rebuild on weight edit) — I'll file it + hold this test, not weaken the assertion. Expected: the DP rebuilds on the etcd-watch snapshot swap (the same path every config-change test relies on), so it should pass.

Verification

tsc --noEmit clean for this file (borrowed node_modules; fresh worktree). Authoritative run is CI's isolated e2e job (etcd + built DP + in-process mock upstreams — no shared-stack contention).

Refs #196 L1, #127 L1.

The sibling weighted-routing-distribution-e2e pins that INITIAL weights
are honored. This closes the L1 gap: a live edit to a weighted model's
weights must propagate through the etcd watch and the weighted
scheduler must REBUILD — a scheduler that cached its weight wheel on
first dispatch and ignored config updates would silently keep serving
the old split.
Deterministic (weight 0 = excluded, per routing-strategies-e2e): start
[wr-edit-a:100, wr-edit-b:0] → assert all dispatches hit A; PUT
/admin/v1/models/:id inverting to [0,100]; the propagation signal is a
virtual probe returning "served by B" (impossible under the old
config); then assert all dispatches hit B. If the scheduler never
rebuilds on a config edit, the post-edit propagation wait times out —
surfacing the regression rather than passing silently.
Refs #196 L1, ai-gateway #127 L1.
@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 8 minutes and 3 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 7a603f8b-22a6-404b-9b39-b12d579b7c2f

📥 Commits

Reviewing files that changed from the base of the PR and between b6a8bd2 and 0053886.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/weighted-routing-edit-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

@moonming

Copy link
Copy Markdown
MemberAuthor

Independent audit (CLAUDE.md §8) — CLEAR

A cold third-party agent reviewed against the actual DP scheduler source. No HIGH/MEDIUM.

  • Weight-0 = exact exclusionweighted_pick (routing.rs) uses strict pick < acc; a 0-weight target is unreachable. Unit-pinned (weighted_pick_zero_weight_target_in_middle_is_never_picked, 2000 trials) + the sibling routing-strategies test. So "all A"/"all B" can't flake on selection.
  • No fall-forward to the 0-weight target — the dispatch loop only advances on a retryable failure; the mock returns 200, which stops dispatch before the in-list 0-weight fallback is contacted. Single dispatch/request, bDelta==0/aDelta==0 guaranteed.
  • PUT body schema-valid + can't 4xxupdate_model requires only display_name; weight: minimum 0; same shape as the create that already validated; assert_unique_name(..., Some(id)) excludes self. Revision is bumped → etcd watch fires.
  • Propagation signal sound — "served by B" is provably impossible under the old [100,0] config; no partial-snapshot window yields it early; the wait fully gates phase 2; baselines snapshotted after the probe so it isn't counted.
  • Catches the regression loudly — a scheduler that ignored the edit → probe stays A → waitConfigPropagation throws (30s). No silent-pass path.
  • Cleanup + isolation — upstreams + app closed; unique etcd prefix/ports/admin-key/display-names/caller-key → safe under maxForks=2 alongside the sibling.

Two LOWs (readiness gate + error-swallowing) are established harness idioms, no change. Verdict: safe to merge.

@moonming
moonming merged commit 5e045e0 into mainJun 5, 2026
7 checks passed
@moonming
moonming deleted the test/issue-L1a-weight-edit-shift branch June 5, 2026 08:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming