.NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

Description

@LoopedBard3

Description

We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

This was found using dotnet/crank and the json scenario tests.

Configuration

  • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
  • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
  • Versions compared:
    • Baseline: .NET 10 (stable)
    • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

Data

CPU Scaling (NET11 on up to 80-core ARM64)

Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

Version-to-version regression (64 cores)

VersionRPSNotes
NET10 (10.0.1)Base RPS
NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

Partial mitigation via environment variables

Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

ASPNETCORE_threadCount=16
DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
DOTNET_EnableWriteXorExecute=0
DOTNET_PerfMapEnabled=1

This suggests the regression is tied to thread count / contention scaling at high core counts.

ASPNET Core KPI version to version regression

The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

Image

Trace Analysis

CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

Total wait/synchronization time: NET10 80.4% → NET11 92.3%

Key changes in NET11:

MethodNET10NET11Δ
WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
PollGCWorker5.52%10.91%+5.39% (doubled)
WaitForSocketEvents0%15.51%+15.51% (was in native)

A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

  • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
  • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
    the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
    harmful to throughput than passive semaphore waits.

Root Cause

The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

Related Issues / PRs

Questions / Ask

  1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
  2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

FYI @DrewScoggins, @VSadov, @jkoritzinsky

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

    Description

    @LoopedBard3

    Description

    We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

    This was found using dotnet/crank and the json scenario tests.

    Configuration

    • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
    • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
    • Versions compared:
      • Baseline: .NET 10 (stable)
      • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

    Data

    CPU Scaling (NET11 on up to 80-core ARM64)

    Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

    Version-to-version regression (64 cores)

    VersionRPSNotes
    NET10 (10.0.1)Base RPS
    NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
    NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
    NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

    Partial mitigation via environment variables

    Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

    ASPNETCORE_threadCount=16
    DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
    DOTNET_EnableWriteXorExecute=0
    DOTNET_PerfMapEnabled=1
    

    This suggests the regression is tied to thread count / contention scaling at high core counts.

    ASPNET Core KPI version to version regression

    The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

    Image

    Trace Analysis

    CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

    Total wait/synchronization time: NET10 80.4% → NET11 92.3%

    Key changes in NET11:

    MethodNET10NET11Δ
    WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
    WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
    LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
    PollGCWorker5.52%10.91%+5.39% (doubled)
    WaitForSocketEvents0%15.51%+15.51% (was in native)

    A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

    • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
    • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
      the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
      harmful to throughput than passive semaphore waits.

    Root Cause

    The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

    PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

    Related Issues / PRs

    Questions / Ask

    1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
    2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

    FYI @DrewScoggins, @VSadov, @jkoritzinsky

    Metadata

    Metadata

    Assignees

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

      Description

      @LoopedBard3

      Description

      We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

      This was found using dotnet/crank and the json scenario tests.

      Configuration

      • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
      • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
      • Versions compared:
        • Baseline: .NET 10 (stable)
        • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

      Data

      CPU Scaling (NET11 on up to 80-core ARM64)

      Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

      Version-to-version regression (64 cores)

      VersionRPSNotes
      NET10 (10.0.1)Base RPS
      NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
      NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
      NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

      Partial mitigation via environment variables

      Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

      ASPNETCORE_threadCount=16
      DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
      DOTNET_EnableWriteXorExecute=0
      DOTNET_PerfMapEnabled=1
      

      This suggests the regression is tied to thread count / contention scaling at high core counts.

      ASPNET Core KPI version to version regression

      The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

      Image

      Trace Analysis

      CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

      Total wait/synchronization time: NET10 80.4% → NET11 92.3%

      Key changes in NET11:

      MethodNET10NET11Δ
      WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
      WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
      LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
      PollGCWorker5.52%10.91%+5.39% (doubled)
      WaitForSocketEvents0%15.51%+15.51% (was in native)

      A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

      • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
      • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
        the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
        harmful to throughput than passive semaphore waits.

      Root Cause

      The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

      PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

      Related Issues / PRs

      Questions / Ask

      1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
      2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

      FYI @DrewScoggins, @VSadov, @jkoritzinsky

      Metadata

      Metadata

      Assignees

      Type

      No type

      Projects

      No projects

        Milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

        Description

        @LoopedBard3

        Description

        We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

        This was found using dotnet/crank and the json scenario tests.

        Configuration

        • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
        • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
        • Versions compared:
          • Baseline: .NET 10 (stable)
          • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

        Data

        CPU Scaling (NET11 on up to 80-core ARM64)

        Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

        Version-to-version regression (64 cores)

        VersionRPSNotes
        NET10 (10.0.1)Base RPS
        NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
        NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
        NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

        Partial mitigation via environment variables

        Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

        ASPNETCORE_threadCount=16
        DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
        DOTNET_EnableWriteXorExecute=0
        DOTNET_PerfMapEnabled=1
        

        This suggests the regression is tied to thread count / contention scaling at high core counts.

        ASPNET Core KPI version to version regression

        The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

        Image

        Trace Analysis

        CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

        Total wait/synchronization time: NET10 80.4% → NET11 92.3%

        Key changes in NET11:

        MethodNET10NET11Δ
        WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
        WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
        LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
        PollGCWorker5.52%10.91%+5.39% (doubled)
        WaitForSocketEvents0%15.51%+15.51% (was in native)

        A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

        • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
        • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
          the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
          harmful to throughput than passive semaphore waits.

        Root Cause

        The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

        PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

        Related Issues / PRs

        Questions / Ask

        1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
        2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

        FYI @DrewScoggins, @VSadov, @jkoritzinsky

        Metadata

        Metadata

        Assignees

        Type

        No type

        Projects

        No projects

          Milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

          Description

          @LoopedBard3

          Description

          We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

          This was found using dotnet/crank and the json scenario tests.

          Configuration

          • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
          • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
          • Versions compared:
            • Baseline: .NET 10 (stable)
            • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

          Data

          CPU Scaling (NET11 on up to 80-core ARM64)

          Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

          Version-to-version regression (64 cores)

          VersionRPSNotes
          NET10 (10.0.1)Base RPS
          NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
          NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
          NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

          Partial mitigation via environment variables

          Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

          ASPNETCORE_threadCount=16
          DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
          DOTNET_EnableWriteXorExecute=0
          DOTNET_PerfMapEnabled=1
          

          This suggests the regression is tied to thread count / contention scaling at high core counts.

          ASPNET Core KPI version to version regression

          The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

          Image

          Trace Analysis

          CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

          Total wait/synchronization time: NET10 80.4% → NET11 92.3%

          Key changes in NET11:

          MethodNET10NET11Δ
          WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
          WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
          LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
          PollGCWorker5.52%10.91%+5.39% (doubled)
          WaitForSocketEvents0%15.51%+15.51% (was in native)

          A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

          • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
          • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
            the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
            harmful to throughput than passive semaphore waits.

          Root Cause

          The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

          PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

          Related Issues / PRs

          Questions / Ask

          1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
          2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

          FYI @DrewScoggins, @VSadov, @jkoritzinsky

          Metadata

          Metadata

          Assignees

          Type

          No type

          Projects

          No projects

            Milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

            Description

            @LoopedBard3

            Description

            We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

            This was found using dotnet/crank and the json scenario tests.

            Configuration

            • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
            • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
            • Versions compared:
              • Baseline: .NET 10 (stable)
              • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

            Data

            CPU Scaling (NET11 on up to 80-core ARM64)

            Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

            Version-to-version regression (64 cores)

            VersionRPSNotes
            NET10 (10.0.1)Base RPS
            NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
            NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
            NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

            Partial mitigation via environment variables

            Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

            ASPNETCORE_threadCount=16
            DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
            DOTNET_EnableWriteXorExecute=0
            DOTNET_PerfMapEnabled=1
            

            This suggests the regression is tied to thread count / contention scaling at high core counts.

            ASPNET Core KPI version to version regression

            The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

            Image

            Trace Analysis

            CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

            Total wait/synchronization time: NET10 80.4% → NET11 92.3%

            Key changes in NET11:

            MethodNET10NET11Δ
            WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
            WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
            LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
            PollGCWorker5.52%10.91%+5.39% (doubled)
            WaitForSocketEvents0%15.51%+15.51% (was in native)

            A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

            • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
            • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
              the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
              harmful to throughput than passive semaphore waits.

            Root Cause

            The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

            PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

            Related Issues / PRs

            Questions / Ask

            1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
            2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

            FYI @DrewScoggins, @VSadov, @jkoritzinsky

            Metadata

            Metadata

            Assignees

            Type

            No type

            Projects

            No projects

              Milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

              Description

              @LoopedBard3

              Description

              We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

              This was found using dotnet/crank and the json scenario tests.

              Configuration

              • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
              • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
              • Versions compared:
                • Baseline: .NET 10 (stable)
                • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

              Data

              CPU Scaling (NET11 on up to 80-core ARM64)

              Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

              Version-to-version regression (64 cores)

              VersionRPSNotes
              NET10 (10.0.1)Base RPS
              NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
              NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
              NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

              Partial mitigation via environment variables

              Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

              ASPNETCORE_threadCount=16
              DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
              DOTNET_EnableWriteXorExecute=0
              DOTNET_PerfMapEnabled=1
              

              This suggests the regression is tied to thread count / contention scaling at high core counts.

              ASPNET Core KPI version to version regression

              The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

              Image

              Trace Analysis

              CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

              Total wait/synchronization time: NET10 80.4% → NET11 92.3%

              Key changes in NET11:

              MethodNET10NET11Δ
              WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
              WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
              LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
              PollGCWorker5.52%10.91%+5.39% (doubled)
              WaitForSocketEvents0%15.51%+15.51% (was in native)

              A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

              • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
              • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
                the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
                harmful to throughput than passive semaphore waits.

              Root Cause

              The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

              PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

              Related Issues / PRs

              Questions / Ask

              1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
              2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

              FYI @DrewScoggins, @VSadov, @jkoritzinsky

              Metadata

              Metadata

              Assignees

              Type

              No type

              Projects

              No projects

                Milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                .NET 11 ASP.NET Core throughput regression on ARM64 at high core counts (16+) (Kestrel JSON benchmark) #127484

                Description

                @LoopedBard3

                Description

                We are observing a 15%+ throughput regression in .NET 11 compared to .NET 10 on the ASP.NET Core Kestrel JSON benchmark, running on ARM64 Linux (Azure Cobalt 100). The regression is most severe at high core counts (>32 cores) where throughput drops dramatically instead of scaling.

                This was found using dotnet/crank and the json scenario tests.

                Configuration

                • Benchmark: ASP.NET Core Kestrel JSON (crank --scenario json --config https://raw.githubusercontent.com/aspnet/Benchmarks/main/scenarios/json.benchmarks.yml <profile information>)
                • Platform: ARM64 Linux (Azure Cobalt 100) (both Ubuntu and Azure Linux 3)
                • Versions compared:
                  • Baseline: .NET 10 (stable)
                  • Regressed: .NET 11.0.0-preview.1.26067.103+bfa3455fa1a8

                Data

                CPU Scaling (NET11 on up to 80-core ARM64)

                Throughput peaks at 32 cores, then drops nearly 50% at 64 cores:

                Version-to-version regression (64 cores)

                VersionRPSNotes
                NET10 (10.0.1)Base RPS
                NET11 alpha (11.0.0-alpha.1.25609.102)-2.5% RPS from baseBefore managed WaitSubsystem (#117788)
                NET11 preview.1 (11.0.0-preview.1.26067.103)-50% RPS from baseAfter WaitSubsystem — -47%
                NET11 preview.1 + thread tuning env vars26% RPS from preview.1Partial recovery — see below

                Partial mitigation via environment variables

                Adding the following environment variables to NET11 preview.1 recovered throughput 26% RPS (still below the base and alpha throughputs):

                ASPNETCORE_threadCount=16
                DOTNET_SYSTEM_NET_SOCKETS_THREAD_COUNT=1
                DOTNET_EnableWriteXorExecute=0
                DOTNET_PerfMapEnabled=1
                

                This suggests the regression is tied to thread count / contention scaling at high core counts.

                ASPNET Core KPI version to version regression

                The regression is also showing up in the KPI dashboard at https://aka.ms/aspnet/benchmarks -> select either Cobalt environment. Current regression is -14.7% throughput for Json Platform test and -14.4% throughput for Json Minimal APIs test.

                Image

                Trace Analysis

                CPU trace comparison (EventPipe SampleProfiler) between .NET 10 and .NET 11 on the same benchmark showed:

                Total wait/synchronization time: NET10 80.4% → NET11 92.3%

                Key changes in NET11:

                MethodNET10NET11Δ
                WaitSubsystem+ThreadWaitInfo.Wait0%4.56%+4.56% (new)
                WaitSubsystem+WaitableObject.Wait_Locked0%3.86%+3.86% (new)
                LowLevelLock.WaitAndAcquire0.27%4.86%+4.59% (18x increase)
                PollGCWorker5.52%10.91%+5.39% (doubled)
                WaitForSocketEvents0%15.51%+15.51% (was in native)

                A comparison with a NET11-alpha build (before the WaitSubsystem change) showed:

                • PollGCWorker doubling was already present in the alpha (10.75%), not caused by Move CoreCLR over to the managed wait subsystem #117788
                • The WaitSubsystem adds ~12.4% new wait CPU but replaces ~12.8% of old LIFO semaphore waits — roughly a wash in total. However,
                  the character changed: LowLevelLock.WaitAndAcquire (active lock contention) quadrupled from 1.21% → 4.86%, which is more
                  harmful to throughput than passive semaphore waits.

                Root Cause

                The primary suspect is PR #117788 ("Move CoreCLR over to the managed wait subsystem"), merged Dec 2025. The managed WaitSubsystem introduces a global process-wide lock (as noted in PR #123921's description) that contends heavily at high core counts.

                PR #123921 ("A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy.") by @VSadov attempted to address this by replacing WaitSubsystem-based blocking with a lightweight portable implementation and adaptive spinning, but was reverted in PR #125193 due to NuGet restore regression. PR #125596 is the pending reapply.

                Related Issues / PRs

                Questions / Ask

                1. Is the high-core-count ARM64 scaling cliff a known dimension of issue [Perf] Linux/x64: 62 Regressions on 1/6/2026 2:06:33 PM +00:00 #123159, or is this a new finding?
                2. Will PR Reapply "A few fixes in the threadpool semaphore. Unify Windows/Unix implementation of LIFO policy." (#125193) #125596 (pending reapply of the threadpool semaphore fixes) address the WaitSubsystem contention seen here?

                FYI @DrewScoggins, @VSadov, @jkoritzinsky

                Metadata

                Metadata

                Assignees

                Type

                No type

                Projects

                No projects

                  Milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions