Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Use SOS to dump managed stack traces from a dump on Windows by jkoritzinsky · Pull Request #82867 · dotnet/runtime · GitHub
Skip to content

Use SOS to dump managed stack traces from a dump on Windows - #82867

Merged
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos
Mar 16, 2023
Merged

Use SOS to dump managed stack traces from a dump on Windows#82867
jkoritzinsky merged 2 commits into
dotnet:mainfrom
jkoritzinsky:sos

Conversation

@jkoritzinsky

Copy link
Copy Markdown
Member

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

@ghost

ghost commented Mar 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.

Issue Details

Download SOS as part of our Helix jobs and load it into cdb to dump managed stack traces when a test crashes or times out.

Also introduces a synthetic test failures to validate behavior (will be removed before merging).

Author:jkoritzinsky
Assignees:jkoritzinsky
Labels:

area-Infrastructure

Milestone:-

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before dotnet-sos existed, we used clrmd directly in RemoteExecutor:
https://github.com/dotnet/arcade/blob/bdc59254cf108e1d48451dc43bb9ebc331cdca7b/src/Microsoft.DotNet.RemoteExecutor/src/RemoteInvokeHandle.cs#L177-L225
We might want to standardize on one or the other.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClrMD there is doing process attach only for timeouts, whereas we're handling both timeouts and crashes here, so it's a little different. Definitely worth considering consolidation though.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There were some talks about standardizing this sort of mechanism (dumping stack traces from a dump to stdout for the build analysis tooling), but with the recent changes to the EngSrv teams, I don't know how much of that work will still happen.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, I understand they're handling different cases. But they can both handle both, so it's overkill to have two different tools used for the same purpose. Simply suggesting we choose one and stick with it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look at the RemoteExecutor implementation and see if we can consolidate something useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, thanks.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to use something like this pattern to get both native and managed stack traces, and we'll likely want to use this model as well to solve #83047 (which focuses more on crashes than timeouts) if we don't move to using dotnet test for the libraries tests in Helix. The RemoteExecutor version has a much better UX around the output though as it has more structure to work with (from CLRMD's APIs).

Comment threadsrc/tests/Common/Coreclr.TestWrapper/CoreclrTestWrapperLib.cs Outdated
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Marking this ready for review as the change works (see the logs for the CoreCLR windows runtime tests), but marked as no-merge as I need to remove the induced test failure.

@hoyosjs for review.

Comment threadsrc/tests/Common/helixpublishwitharcade.proj Outdated
@hoyosjs

Copy link
Copy Markdown
Member

@jkoritzinsky this looks good now. It's not loading the PDBs for managed for for line numbers, but that's separate.

@jkoritzinskyjkoritzinsky removed the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Mar 16, 2023
@jkoritzinsky

Copy link
Copy Markdown
MemberAuthor

Networking failures are known issues, wasm timeout is unrelated. Merging this in. I'll file an issue to follow-up on providing a consistent story for dumping stacks on crashes.

@jkoritzinsky
jkoritzinsky merged commit ef43a7b into dotnet:mainMar 16, 2023
@jkoritzinsky
jkoritzinsky deleted the sos branch March 16, 2023 20:14
@ghostghost locked as resolved and limited conversation to collaborators Apr 16, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jkoritzinsky@hoyosjs@stephentoub@danmoseley