Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

Description

@Astro-Han

e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

Evidence

Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

RunBranch
32741751712main (push)
32737500876fix/align-usage-request-counts
32735909741feat/runtime-host-update-reconciliation
32731153705Connect-Custom-relay-fetch-models
32722992191feat/github-copilot-device-flow-login

Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

Why the assertion is wider than its intent

The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

  1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
  2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

What not to do

apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

Suggested direction

Scope the observation to what the test actually means:

  • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
  • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
  • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

简体中文

e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

证据

在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

运行分支
32741751712main(push)
32737500876fix/align-usage-request-counts
32735909741feat/runtime-host-update-reconciliation
32731153705Connect-Custom-relay-fetch-models
32722992191feat/github-copilot-device-flow-login

每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

为什么这个断言比它的意图更宽

该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

  1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
  2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

不要做什么

apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

建议方向

把观察范围收敛到测试真正想表达的意思:

  • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
  • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
  • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

Metadata

Metadata

Labels

bugSomething isn't workinghelp wantedExtra attention is needed

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

    Description

    @Astro-Han

    e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

    Evidence

    Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

    RunBranch
    32741751712main (push)
    32737500876fix/align-usage-request-counts
    32735909741feat/runtime-host-update-reconciliation
    32731153705Connect-Custom-relay-fetch-models
    32722992191feat/github-copilot-device-flow-login

    Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

    Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

    Why the assertion is wider than its intent

    The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

    1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
    2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

    So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

    What not to do

    apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

    Suggested direction

    Scope the observation to what the test actually means:

    • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
    • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
    • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

    Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

    简体中文

    e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

    证据

    在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

    运行分支
    32741751712main(push)
    32737500876fix/align-usage-request-counts
    32735909741feat/runtime-host-update-reconciliation
    32731153705Connect-Custom-relay-fetch-models
    32722992191feat/github-copilot-device-flow-login

    每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

    请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

    为什么这个断言比它的意图更宽

    该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

    1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
    2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

    所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

    不要做什么

    apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

    建议方向

    把观察范围收敛到测试真正想表达的意思:

    • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
    • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
    • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

    接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

    Metadata

    Metadata

    Labels

    bugSomething isn't workinghelp wantedExtra attention is needed

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

      Description

      @Astro-Han

      e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

      Evidence

      Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

      RunBranch
      32741751712main (push)
      32737500876fix/align-usage-request-counts
      32735909741feat/runtime-host-update-reconciliation
      32731153705Connect-Custom-relay-fetch-models
      32722992191feat/github-copilot-device-flow-login

      Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

      Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

      Why the assertion is wider than its intent

      The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

      1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
      2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

      So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

      What not to do

      apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

      Suggested direction

      Scope the observation to what the test actually means:

      • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
      • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
      • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

      Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

      简体中文

      e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

      证据

      在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

      运行分支
      32741751712main(push)
      32737500876fix/align-usage-request-counts
      32735909741feat/runtime-host-update-reconciliation
      32731153705Connect-Custom-relay-fetch-models
      32722992191feat/github-copilot-device-flow-login

      每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

      请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

      为什么这个断言比它的意图更宽

      该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

      1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
      2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

      所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

      不要做什么

      apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

      建议方向

      把观察范围收敛到测试真正想表达的意思:

      • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
      • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
      • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

      接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

      Metadata

      Metadata

      Labels

      bugSomething isn't workinghelp wantedExtra attention is needed

      Type

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

        Description

        @Astro-Han

        e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

        Evidence

        Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

        RunBranch
        32741751712main (push)
        32737500876fix/align-usage-request-counts
        32735909741feat/runtime-host-update-reconciliation
        32731153705Connect-Custom-relay-fetch-models
        32722992191feat/github-copilot-device-flow-login

        Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

        Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

        Why the assertion is wider than its intent

        The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

        1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
        2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

        So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

        What not to do

        apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

        Suggested direction

        Scope the observation to what the test actually means:

        • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
        • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
        • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

        Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

        简体中文

        e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

        证据

        在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

        运行分支
        32741751712main(push)
        32737500876fix/align-usage-request-counts
        32735909741feat/runtime-host-update-reconciliation
        32731153705Connect-Custom-relay-fetch-models
        32722992191feat/github-copilot-device-flow-login

        每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

        请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

        为什么这个断言比它的意图更宽

        该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

        1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
        2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

        所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

        不要做什么

        apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

        建议方向

        把观察范围收敛到测试真正想表达的意思:

        • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
        • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
        • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

        接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

        Metadata

        Metadata

        Labels

        bugSomething isn't workinghelp wantedExtra attention is needed

        Type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

          Description

          @Astro-Han

          e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

          Evidence

          Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

          RunBranch
          32741751712main (push)
          32737500876fix/align-usage-request-counts
          32735909741feat/runtime-host-update-reconciliation
          32731153705Connect-Custom-relay-fetch-models
          32722992191feat/github-copilot-device-flow-login

          Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

          Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

          Why the assertion is wider than its intent

          The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

          1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
          2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

          So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

          What not to do

          apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

          Suggested direction

          Scope the observation to what the test actually means:

          • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
          • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
          • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

          Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

          简体中文

          e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

          证据

          在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

          运行分支
          32741751712main(push)
          32737500876fix/align-usage-request-counts
          32735909741feat/runtime-host-update-reconciliation
          32731153705Connect-Custom-relay-fetch-models
          32722992191feat/github-copilot-device-flow-login

          每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

          请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

          为什么这个断言比它的意图更宽

          该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

          1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
          2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

          所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

          不要做什么

          apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

          建议方向

          把观察范围收敛到测试真正想表达的意思:

          • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
          • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
          • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

          接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

          Metadata

          Metadata

          Labels

          bugSomething isn't workinghelp wantedExtra attention is needed

          Type

          Projects

          No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

            Description

            @Astro-Han

            e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

            Evidence

            Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

            RunBranch
            32741751712main (push)
            32737500876fix/align-usage-request-counts
            32735909741feat/runtime-host-update-reconciliation
            32731153705Connect-Custom-relay-fetch-models
            32722992191feat/github-copilot-device-flow-login

            Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

            Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

            Why the assertion is wider than its intent

            The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

            1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
            2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

            So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

            What not to do

            apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

            Suggested direction

            Scope the observation to what the test actually means:

            • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
            • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
            • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

            Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

            简体中文

            e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

            证据

            在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

            运行分支
            32741751712main(push)
            32737500876fix/align-usage-request-counts
            32735909741feat/runtime-host-update-reconciliation
            32731153705Connect-Custom-relay-fetch-models
            32722992191feat/github-copilot-device-flow-login

            每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

            请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

            为什么这个断言比它的意图更宽

            该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

            1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
            2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

            所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

            不要做什么

            apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

            建议方向

            把观察范围收敛到测试真正想表达的意思:

            • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
            • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
            • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

            接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

            Metadata

            Metadata

            Labels

            bugSomething isn't workinghelp wantedExtra attention is needed

            Type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

              Description

              @Astro-Han

              e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

              Evidence

              Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

              RunBranch
              32741751712main (push)
              32737500876fix/align-usage-request-counts
              32735909741feat/runtime-host-update-reconciliation
              32731153705Connect-Custom-relay-fetch-models
              32722992191feat/github-copilot-device-flow-login

              Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

              Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

              Why the assertion is wider than its intent

              The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

              1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
              2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

              So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

              What not to do

              apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

              Suggested direction

              Scope the observation to what the test actually means:

              • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
              • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
              • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

              Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

              简体中文

              e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

              证据

              在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

              运行分支
              32741751712main(push)
              32737500876fix/align-usage-request-counts
              32735909741feat/runtime-host-update-reconciliation
              32731153705Connect-Custom-relay-fetch-models
              32722992191feat/github-copilot-device-flow-login

              每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

              请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

              为什么这个断言比它的意图更宽

              该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

              1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
              2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

              所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

              不要做什么

              apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

              建议方向

              把观察范围收敛到测试真正想表达的意思:

              • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
              • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
              • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

              接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

              Metadata

              Metadata

              Labels

              bugSomething isn't workinghelp wantedExtra attention is needed

              Type

              Projects

              No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                Flaky test: slash-command-menu "keeps its container and skills group across projection refreshes" #3727

                Description

                @Astro-Han

                e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes fails intermittently on unrelated branches, and it is currently the single most frequent cause of a red test check.

                Evidence

                Across the 60 most recent failed CI runs, Desktop e2e was the failing step in 9. Five of those were this one test, on five different branches:

                RunBranch
                32741751712main (push)
                32737500876fix/align-usage-request-counts
                32735909741feat/runtime-host-update-reconciliation
                32731153705Connect-Custom-relay-fetch-models
                32722992191feat/github-copilot-device-flow-login

                Every one reports 1 failed, 61 passed, and every one fails at slash-command-menu.spec.ts:237 — the expect.poll(...).toBe(0) on the mutation counter.

                Note the branches have nothing in common; the main occurrence is the same failure on a commit whose own PR head was green.

                Why the assertion is wider than its intent

                The test arms a MutationObserver before three same-content projection refreshes and asserts that no [role="listbox"] or [role="group"] node was removed. That is the right property to guard (#2667). Two details make it report more than that property:

                1. The counter only increases. Once removals is non-zero it can never return to 0, so expect.poll(..., { timeout: 3_000 }) cannot retry into success — it just waits 3 seconds and then fails. The poll adds latency, not tolerance.
                2. The observation window is not scoped to the refreshes. The observer runs on document.body with subtree: true for the whole poll window, and it counts any removed element that contains a listbox or group anywhere beneath it. A toast, tooltip, or popup unmounting during those 3 seconds is counted the same as the regression the test is guarding.

                So a failure here means "something containing a group was removed somewhere in the document during a 3-second window", which is a superset of "the skills group was torn down by the projection refresh".

                What not to do

                apps/desktop/e2e/playwright.config.ts sets retries: 0 with an explicit rationale — flakes should fail loudly. That is a deliberate choice and this issue is not a request to change it. Adding retries would hide this rather than fix it.

                Suggested direction

                Scope the observation to what the test actually means:

                • Start the observer immediately before the refresh rounds and disconnect it as soon as they complete, instead of leaving it running for the poll window.
                • Observe the menu container rather than document.body, so unrelated overlays cannot contribute.
                • Assert the count once after the refreshes have settled, rather than polling a monotonically increasing counter.

                Anyone picking this up should confirm the failure is genuinely the observation window and not a real intermittent teardown — the trace and video are retained on failure (trace: 'retain-on-failure'), and the artifacts on the runs above are the place to start.

                简体中文

                e2e/slash-command-menu.spec.ts:166 › an open menu keeps its container and skills group across projection refreshes 会在互不相干的分支上间歇性失败,目前是导致 test 检查变红的最主要单一原因

                证据

                在最近 60 次失败的 CI 运行中,Desktop e2e 是失败步骤的有 9 次。其中 5 次是同一个用例,分布在五个不同分支上:

                运行分支
                32741751712main(push)
                32737500876fix/align-usage-request-counts
                32735909741feat/runtime-host-update-reconciliation
                32731153705Connect-Custom-relay-fetch-models
                32722992191feat/github-copilot-device-flow-login

                每一次都是 1 failed, 61 passed,且每一次都失败在 slash-command-menu.spec.ts:237——也就是对 mutation 计数器的 expect.poll(...).toBe(0)

                请注意这些分支之间毫无共同点;main 上那次是同一个失败,而它对应的 PR head 自身是绿的。

                为什么这个断言比它的意图更宽

                该测试在三次同内容的 projection 刷新之前装上一个 MutationObserver,断言没有任何 [role="listbox"][role="group"] 节点被移除。要守的性质是对的(#2667)。但有两处细节让它报出的东西超出了这个性质:

                1. 计数器只增不减。removals 一旦非零就再也回不到 0,因此 expect.poll(..., { timeout: 3_000 }) 不可能重试成真——它只是等满 3 秒然后失败。这个 poll 带来的是延迟,不是容忍度。
                2. 观察窗口没有限定在刷新期间。 观察器以 subtree: true 挂在 document.body 上、贯穿整个 poll 窗口,并且只要被移除的元素内部任何位置含有 listbox 或 group 就计数。在那 3 秒里卸载的一个 toast、tooltip 或弹层,与测试真正要守的那个回归被同等对待。

                所以这里的失败意味着"在 3 秒窗口内,文档中某处移除了一个含 group 的元素"——这是"skills group 被 projection 刷新拆掉"的一个超集

                不要做什么

                apps/desktop/e2e/playwright.config.ts 设的是 retries: 0,并写明了理由——flake 应当响亮地失败。那是有意的选择,本 issue 不是要求改它。加重试是把问题盖住,不是修好。

                建议方向

                把观察范围收敛到测试真正想表达的意思:

                • 在刷新轮次开始前才装上观察器,刷新一结束就 disconnect,而不是让它在整个 poll 窗口里一直跑。
                • 观察菜单容器本身而不是 document.body,这样无关的浮层无法贡献计数。
                • 在刷新稳定之后一次性断言计数,而不是去轮询一个单调递增的计数器。

                接手的人应当先确认这确实是观察窗口的问题,而不是一次真实的间歇性拆除——失败时 trace 与 video 都会保留(trace: 'retain-on-failure'),上面几次运行的产物就是起点。

                Metadata

                Metadata

                Labels

                bugSomething isn't workinghelp wantedExtra attention is needed

                Type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions