Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Fix some scaling issues with the global queue in the thread pool - #69386

Merged
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit
Jun 7, 2022
Merged

Fix some scaling issues with the global queue in the thread pool#69386
kouvel merged 3 commits into
dotnet:mainfrom
kouvel:TpSplit

Conversation

@kouvel

@kouvelkouvel commented May 16, 2022

Copy link
Copy Markdown
Contributor
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue

Fixes#67845

- The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
- I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
- Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
- Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
- When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
- In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
@kouvelkouvel added this to the 7.0.0 milestone May 16, 2022
@kouvelkouvel self-assigned this May 16, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details
  • The global concurrent queue is not scaling well in some situations on machines with a large number of processors. A lot of contention is seen in dequeuing work on some benchmarks.
  • I initially tried switching to queuing work from thread pool threads more independently with more efficient work stealing, but it looks like that may need more investigation/experimentation. This is a simpler change that seems to work reasonably well for now.
  • Beyond 32 procs, added additional concurrent queues, one per 16 procs. When a worker thread begins dispatching work, it assigns one of those additional queues to itself, trying to limit assignments to 16 worker threads per queue.
  • Work items queued by a worker thread queues to the assigned queue. The worker thread dequeues from the assigned queue first, and later tries to dequeue from other queues. Work items queued by non-thread-pool threads continue to go into the global queue.
  • When a worker thread stops dispatching work items, it unassigns itself from the queue, and may transfer work items from it if it was the last thread assigned to the queue.
  • In the observed cases, work items are distributed to the different queues and contention is reduced due to a limited number of threads operating on each queue
Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:7.0.0

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

RPS numbers for ASP.NET perf on an 80-proc arm64 Linux machine below. For this machine's comparison I included dotnet/aspnetcore#40476 on both the before and after sides because in some cases the thread pool global queue bottleneck was only showing up after that fix.

RPSBeforeAfterDiff
PlaintextPlatform78964861183016449.8%
JsonPlatform3109501146701268.8%
FortunesPlatform33495543797630.8%
Plaintext9785381101873654.1%
Json2700471089592303.5%
Fortunes24493733105435.2%

No significant change on a 48-proc x64 Linux machine:

RPSBeforeAfterDiff
PlaintextPlatform12042603120515970.1%
JsonPlatform13521481348159-0.3%
FortunesPlatform4707024786761.7%
Plaintext776030578351561.0%
Json105309610594310.6%
Fortunes3976884000520.6%

No significant change on other citrine machines where only one queue is used.

@kouvel
kouvel requested review from janvorli and mangod9May 16, 2022 09:58
@kunalspathak

Copy link
Copy Markdown
Contributor

FYI - @sebastienros@JulieLeeMSFT

@sebastienros

Copy link
Copy Markdown
Member

And these numbers are without solving the same issue in aspnet? In my tests this was still required so I find this very promising.

@JulieLeeMSFT

Copy link
Copy Markdown
Member

No significant change on a 48-proc x64 Linux machine:

@kouvel is this RPS measurement?

@kouvel

kouvel commented May 16, 2022

Copy link
Copy Markdown
ContributorAuthor

And these numbers are without solving the same issue in aspnet?

For the ampere machine I included your fix from dotnet/aspnetcore#40476 in the before and after columns so the numbers include the aspnet fix, as in some tests the thread pool scaling issue was only showing up after that. For the other machines I didn't include the aspnet fix.

@kouvel is this RPS measurement?

Yes, these are RPS numbers. I'll check the mean latency on the ampere machine and will add those.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Here are the mean latency numbers for ASP.NET perf on the 80-proc arm64 Linux machine:

LatencyBefore (ms)After (ms)Diff (ms)
PlaintextPlatform1.431.640.21
JsonPlatform1.860.56-1.30
FortunesPlatform1.861.61-0.25
Plaintext0.330.31-0.02
Json0.970.31-0.66
Fortunes1.521.18-0.34

@kunalspathak

Copy link
Copy Markdown
Contributor

@kouvel - any update on this PR? Is there anything left?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Needs a review, @janvorli or @mangod9 could you please take a look?

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with one nit. I am sorry it took me so long, I've forgotten about this PR.

@stephentoub

Copy link
Copy Markdown
Member

This maintains the FIFO nature of work items queued by a given thread to the global queue. Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Technically it loses the FIFO nature of multiple threads queueing to the global queue, right? e.g. if thread 1 queues a work item, then signals to thread 2 which in turn queues a work item, it's possible thread 2's work item could be dequeued before thread 1's, yes? I assume we think the chances of that negatively impacting something are slim.

Yea it's a tradeoff. Maintaining a strict global FIFO order for work items comes with scalability issues. Based on my other testing, breaking the FIFO order more (such as with independent queuing of work by threads with work stealing), currently has some issues with ordering of certain work items, although that model may work better eventually. This change tries to maintain some FIFO ordering while breaking some of it to scale better. I don't anticipate many issues from it since where it's applicable, the scalability issue seems to be a larger issue. Another potential tradeoff is that in some cases on larger machines with low load, some threads may have to spend more CPU time searching more queues for work, resulting in more CPU time for the same work done. I haven't seen it yet, though that one may be workable if it becomes an issue.

@kouvel
kouvel merged commit deb4044 into dotnet:mainJun 7, 2022
@kouvel
kouvel deleted the TpSplit branch June 7, 2022 15:01
kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 7, 2022
* Update SOS to show thread pool work items from new queues
- Depends on dotnet/runtime#69386
- The PR above added new queues of work items. This change updates the `ThreadPool -wi` command to include showing work items from those new queues.
- Previously on x86 it looks like it was reading 8 bytes from an array element of pointer size and only using the lower 4 bytes. Fixed to read only pointer size in a few places.
@EgorBo

EgorBo commented Jun 16, 2022

Copy link
Copy Markdown
Member

@adamsitnik

Copy link
Copy Markdown
Member

@kouvel is there any chance you are going to blog about it? or do a Platform talk? it looks like a very interesting problem, I would love to hear the whole story behind it

@kouvel

Copy link
Copy Markdown
ContributorAuthor

There's some more info about the issue I ran into in dotnet/aspnetcore#41391. There are more things to investigate, including polling for IO on worker threads. It is to be determined what other changes may help and if it would work better than currently. We can chat if you'd like.

@ghostghost locked as resolved and limited conversation to collaborators Aug 5, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The thread pool's global queue doesn't scale well on machines with a large processor count

9 participants

@kouvel@kunalspathak@sebastienros@JulieLeeMSFT@stephentoub@EgorBo@adamsitnik@janvorli@mangod9