Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Improve the rate of thread injection for blocking due to sync-over-async - #53471

Merged
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait
Jun 9, 2021
Merged

Improve the rate of thread injection for blocking due to sync-over-async#53471
kouvel merged 12 commits into
dotnet:mainfrom
kouvel:TpTaskWait

Conversation

@kouvel

@kouvelkouvel commented May 30, 2021

Copy link
Copy Markdown
Contributor
  • FixesImprove the rate of thread injection for blocking due to sync-over-async #52558
  • Some miscellaneous changes:
    • _minThreads and _maxThreads were being modified inside their own lock and used inside the hill climbing lock, so it made sense to merge the two locks
    • Separated NumThreadsGoal from ThreadCounts into its own field to simplify some code. The goal is an estimated target and doesn't need to be in perfect sync with the other values in ThreadCounts. The goal was already only modified inside a lock.
    • Removed some unnecessary volatile accesses to simplify

@kouvelkouvel added this to the 6.0.0 milestone May 30, 2021
@kouvelkouvel self-assigned this May 30, 2021
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes #52558

Author:kouvel
Assignees:kouvel
Labels:

area-System.Threading

Milestone:6.0.0

@kouvel

kouvel commented May 30, 2021

Copy link
Copy Markdown
ContributorAuthor
  • Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
  • Verified config vars are working as expected
  • Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
  • Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
  • Checked the relevant cases from https://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
  • The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

@kouvel
kouvel requested a review from marek-safar as a code ownerMay 30, 2021 02:14
@davidfowl

davidfowl commented May 30, 2021

Copy link
Copy Markdown
Member

What do you think about doing this in monitor.wait as well?

@kouvel

Copy link
Copy Markdown
ContributorAuthor

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

@kouvelkouvelJun 2, 2021

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some of the criteria used:

  • Have a good replacement for setting the MinThreads as a workaround
    • This would now be to set ThreadsToAddWithoutDelay to the equivalent and set MaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
    • MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
  • Use progressive delays to avoid creating too many threads too quickly
    • Without that, it would be conceivable that MaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
    • The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
    • The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
    • Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
    • It's not always clear that adding more threads would help, more so in starvation-type cases
  • No one set of defaults will work well for all cases, use conservative defaults to start with
    • The current defaults are much more agressive than before
    • The MaxThreadsToAddBeforeFallback_ProcCountFactor of 10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
    • The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
    • The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
  • Make things sufficiently configurable
    • It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
    • Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
    • I expect I would suggest most users running into bad blocking due to sync-over-async to configure ThreadsToAddWithoutDelay and MaxThreadsToAddBeforeFallback, and perhaps MaxDelayUntilFallbackMs to control the thread injection latency for spikes
  • I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
    • Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Decided to remove MaxThreadsToAddBeforeFallback and renamed MaxDelayBeforeFallbackMs to MaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

Rebased to fix conflicts

@kouvel

Copy link
Copy Markdown
ContributorAuthor

I believe I have addressed the feedback shared so far, any other feedback?

@mangod9mangod9 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Also would be good to create a doc issue for the new configs.

@kouvel

Copy link
Copy Markdown
ContributorAuthor

LGTM, assuming that some simple scenarios which were exhibiting deadlocks are now more responsive, and there don't seem to be any other regressions?

Thanks! Yes some simple scenarios involving sync-over-async are much more responsive by default, and can be configured sufficiently well to work around high levels of blocking if necessary. I haven't seen any regressions in what I tested above.

Filed dotnet/docs#24566 for updating docs.

@kouvel
kouvel merged commit a7a2fd6 into dotnet:mainJun 9, 2021
@kouvel
kouvel deleted the TpTaskWait branch June 9, 2021 18:13
@davidfowl

Copy link
Copy Markdown
Member

Looking forward to this!

kouvel added a commit to dotnet/diagnostics that referenced this pull request Jun 10, 2021
- Depends on dotnet/runtime#53471
- In the above change `ThreadCounts._data` was changed from `ulong` to `uint`. Its usage in SOS was reading 8 bytes and only using the lower 4 bytes. Updated to read 4 bytes instead. No functional change, just updated to match.
@ghostghost locked as resolved and limited conversation to collaborators Jul 10, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve the rate of thread injection for blocking due to sync-over-async

6 participants

@kouvel@davidfowl@benaadams@stephentoub@janvorli@mangod9