Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Reduce contention for Datacontract Serialization by Daniel-Svensson · Pull Request #70668 · dotnet/runtime · GitHub
Skip to content

Reduce contention for Datacontract Serialization - #70668

Closed
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention
Closed

Reduce contention for Datacontract Serialization#70668
Daniel-Svensson wants to merge 7 commits into
dotnet:mainfrom
Daniel-Svensson:datacontract_contention

Conversation

@Daniel-Svensson

@Daniel-SvenssonDaniel-Svensson commented Jun 13, 2022

Copy link
Copy Markdown
Contributor

Making profiling/stress testing my rewrite of OpenRiaServices from WCF to aspnet core I identified a bottleneck in the datacontract serializer.

I detected some contention for an API returning around 20 000 objects of a fairly simple type with a datacontract surrogate which looks the same.

This fix is made in 2 steps,

  • Improve GetBuiltInDataContract performance
  • Improve GetID performance

Loadtesting was done from local machine so the load testing (Netling/Westwind Web Surge) also consume resources

Initially GetBuiltInDataContract was the source of contention.
Contention time for 20s run:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary instead removing the contention.

When fixing that most contention seemed to move to GetId:
Contention time for 20s run after first commit:
image

  • The lock + Dictionary was replaced by a ConcurrentDictionary and the lock for the cache is only taken on update

Other possible solutions / comments

  • The first fix maybe setup the whole dictionary on first access and then keeping the dictionary read only ?
    I decided to keep current logic in order to keep the change small and not affect startup/first access performance to much,
  • I looked into caching a static Func for the factory method, but I got problem due to the RequiresUnreferencedCode so i used a static lambda instead (which seems to be cached by roslyn, but with an extra if statement)

Numbers

On 16 AMD 5800X

8 thead loading

net7 preview 4net7 preview 4 + first commitnet7 preview 4 + second commitsnet7 preview 4 + 3rd commits
req/sec115 req/sec235 req/sec255 req/sec258 req/sec
data served5.1gb10.5gb11,3gb11,5gb
95%th82,7ms47,1ms44,7ms42,43ms
99%th94,9ms72ms63,6ms51,51ms

16 threads loading

From left to right:

  • net7 preview 4 "original" (with r2r)
  • net7 preview 4 + first commit (~80% cpu load, including load generation)
  • net7 preview 4 + second commits (100% cpu load including load generation)
    image
  • net7 preview 4 + 3rd commit (lazy) (100% cpu load)

* net7 preview 4 + 4th commit (lazy + changed key) *

1 or 2 threads loaded

No significant difference between the versions to draw any conclusions.
update August: performance seems to be better even with no or little concurrency.

  • ~5% for 1 thread (78 -> 82 rps)
  • ~16% for 2 threads (135->157 rps)

Other

  • When only serializing 10 objects per query the improvement drops there is still a measureable improvement ~165 000 rps vs ~<143 000 rps for 16 threads (console logging set to warning)

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-Serialization and removed community-contribution Indicates that the PR has been added by a community member labels Jun 13, 2022
@Daniel-Svensson

Daniel-Svensson commented Jun 14, 2022

Copy link
Copy Markdown
ContributorAuthor

Commenting reply here as well since it is hidden when resolved.

I did a small benchmarks (single threaded), which just invokes GetId for different keys with almost no adds just to measure overhead for the non-contented case and settled for using concurrent dictionary + lazy.

BenchmarkDotNet=v0.13.1, OS=Windows 10.0.22000
AMD Ryzen 7 5800X, 1 CPU, 16 logical and 8 physical cores
.NET SDK=7.0.100-preview.4.22252.9
[Host] : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
DefaultJob : .NET 6.0.5 (6.0.522.21309), X64 RyuJIT
MethodMeanErrorStdDev
Original2.693 μs0.0211 μs0.0197 μs
Id2 (using collectionmarshal)2.853 μs0.0160 μs0.0149 μs
Lazy3.435 μs0.0252 μs0.0236 μs
LazyAndValueKey2.565 μs0.0119 μs0.0111 μs
LazyAndValueKeyWithEqualsComparer2.558 μs0.0089 μs0.0083 μs
Faulty3.407 μs0.0238 μs0.0211 μs
RwlockSlim5.513 μs0.0208 μs0.0194 μs
HashSet2.784 μs0.0067 μs0.0056 μs
ConcurrentDictionary<, int> (commit 7)1.713 μs0.0157 μs0.0147 μs

Update: I switched to use RuntimeTypeHandle as key in order to make the code simpler (since there is no ThreadStatic storage required to cache type handles) which also had the benefit of better performance than original code.
It you think that is not a good solution then feel free to skip the last commit.

In original code GetBuiltInDataContract took 9,8% of time GetId took 0.5% of time.
After commit 3 the numbers where numbers are 0,3% and 0,08%

@StephenMolloy

Copy link
Copy Markdown
Member

We have a very large PR in the works trying to reconcile the strange port of DCS that has been in .Net Core with the more complete version in .Net 4.8. (#71752). In doing this reconciliation, I believe this issue is also addressed - in a similar but not exactly the same manner. After that PR goes through though, you should see performance improvement here. If not, please reopen the issue.

@Daniel-Svensson

Copy link
Copy Markdown
ContributorAuthor

@StephenMolloy do you have time frame for your large pr, Is there any chance of this getting reviewed in time to make it possible for a merge before the net7 cut-off date at August 15 ?

I did a quick test if your pr and it only improves performance marginally so I think it still makes sense to improve the 2 methods that are changed here.
FB924B45071F42D59E3C6E2B51073905

@StephenMolloy

Copy link
Copy Markdown
Member

As mentioned earlier, a similar approach on this particular issue has been brought back from .Net 4.8 in our mega reconciliation of 7.0 with 4.8. #71752 has been checked in now. Hopefully you see this show up in your performance testing of RC1.

@ghostghost locked as resolved and limited conversation to collaborators Sep 12, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Daniel-Svensson@StephenMolloy@drieseng@ts-223