Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add reproducible MSTest performance benchmarks - #10843

Merged
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance
Aug 28, 2026
Merged

Add reproducible MSTest performance benchmarks#10843
Amaury Levé (Evangelink) merged 8 commits into
microsoft:mainfrom
Evangelink:dev/amauryleve/compare-test-framework-performance

Conversation

@Evangelink

@EvangelinkAmaury Levé (Evangelink) commented Aug 28, 2026

Copy link
Copy Markdown
Member

MSTest has several performance-sensitive execution modes, but the existing profiler runner did not produce statistically useful, outcome-validated results or directly compare MTP and VSTest. This adds a repeatable benchmark foundation for tracking those costs without turning noisy hosted-agent timings into a pull request gate.

Related to #9312 and #10549.

What changed

  • Extend the end-to-end runner with one warmup and five measured runs, expected test-count validation, median/IQR output, runtime metadata, and stable JSON artifacts.
  • Add comparable MTP standalone, MTP dotnet test, VSTest, Rooting, ReflectionFree, and NativeAOT lanes across the existing plain, data-driven, and lifecycle workloads.
  • Add BenchmarkDotNet allocation coverage for MSTest-to-MTP node conversion.
  • Preserve partial results when one nightly lane fails, reset incompatible historical baselines, and compare only matching OS, CI image, architecture, processor count, effective worker count, runner runtime, target framework, and configuration dimensions.
  • Collect Windows/Linux artifacts plus Linux microbenchmarks without gating pull requests.

CPU and working-set metrics are intentionally omitted for dotnet test lanes because the launched CLI process does not include its test-host children. Wall-clock timings remain available, and the README documents which dimensions must match before comparing runs.

Validation

  • Full repository build through build.cmd
  • Debug performance-runner build with zero warnings
  • Release benchmark and MSTest.slnf builds
  • MTP and VSTest dotnet test lanes with 10,000/10,000 passing tests
  • Reflection, Rooting, ReflectionFree, and NativeAOT standalone lanes with 10,000/10,000 passing tests
  • BenchmarkDotNet ShortRun for all node-conversion benchmarks
  • Historical and new current-report parsing, schema-v3 baseline matching, and environment/worker-isolation behavior

Add acceptance coverage that verifies the published assembly contains a ReadyToRun header and executes its tests successfully.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Add validated end-to-end MTP, VSTest, source-generation, and NativeAOT benchmark lanes plus BenchmarkDotNet allocation coverage and non-gating nightly collection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
Keep VSTest benchmark artifacts nullable while preserving a clear runtime guard for profiler pipelines that launch the MTP executable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI balanced review requested due to automatic review settings August 28, 2026 14:05

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 4 Medium severity · 1 Low severity

New issues introduced by this change (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The…
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.…
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to…
What changed in this PR

Adds reproducible end-to-end and allocation benchmarks for MSTest, with nightly artifact collection and compatibility with rolling performance reports.

Changes:

  • Adds validated multi-run MTP, VSTest, source-generation, and NativeAOT benchmarks.
  • Adds BenchmarkDotNet node-conversion measurements and ReadyToRun coverage.
  • Extends nightly collection, reporting, and partial-result handling.
FileDescription
TestFx.slnxIncludes benchmark project.
MSTest.slnfIncludes benchmark project.
Directory.Packages.propsAdds BenchmarkDotNet version.
test/​Performance/​README.mdDocuments benchmark usage.
test/​Performance/​MSTest.Performance.Runner/​Program.csDefines expanded benchmark matrix.
test/​Performance/​MSTest.Performance.Runner/​Pipeline.csAdds context cleanup.
test/​Performance/​MSTest.Performance.Runner/​PipelinesRunner.csContinues after lane failures.
test/​Performance/​MSTest.Performance.Runner/​Context.csMakes disposal repeat-safe.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.csAdds platform and source-generation modes.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario2.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario3.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario4.csAdds expected-count metadata.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.csImplements measurement and JSON reporting.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PlainProcess.csUses repeatable standalone measurements.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTestProcess.csMeasures MTP and VSTest CLI execution.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.csPublishes NativeAOT assets.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetMuxer.csSupports nullable standalone hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetTrace.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​PerfviewRunner.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​VSDiagnostics.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Runner/​Steps/​ConcurrencyVisualizer.csUses validated test hosts.
test/​Performance/​MSTest.Performance.Benchmarks/​Program.csAdds BenchmarkDotNet entry point.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTestTestNodeConverterBenchmarks.csBenchmarks node conversion allocations.
test/​Performance/​MSTest.Performance.Benchmarks/​MSTest.Performance.Benchmarks.csprojConfigures benchmark executable.
test/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.csAdds ReadyToRun publish test.
src/​Adapter/​MSTest.TestAdapter/​MSTest.TestAdapter.csprojGrants benchmark internal access.
src/​Adapter/​MSTestAdapter.PlatformServices/​MSTestAdapter.PlatformServices.csprojGrants benchmark internal access.
.github/​workflows/​perf-timing-nightly.ymlExpands nightly collection.
.github/​scripts/​compare_perf_timings.pySupports new and historical schemas.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadtest/Performance/MSTest.Performance.Runner/Scenarios/Scenario1.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs Outdated
Comment thread.github/scripts/compare_perf_timings.py
Comment threadtest/IntegrationTests/MSTest.Acceptance.IntegrationTests/ReadyToRunTests.cs Outdated
Comment threadtest/Performance/MSTest.Performance.Runner/Program.cs
Key rolling baselines by environment and configuration, report the requested target framework, distinguish runner runtime metadata, and remove duplicate ReadyToRun coverage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:24
@EvangelinkAmaury Levé (Evangelink) added the state/needs-review Awaiting review from the team. label Aug 28, 2026

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity

New issues introduced by this change (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is…
Issues resolved since last review (5)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Program.cs — This existing pipeline keeps the same baseline key while changing its binaries from Debug to… View resolved comment
Low severitytest/​IntegrationTests/​MSTest.Acceptance.IntegrationTests/​ReadyToRunTests.cs — This adds a second expensive ReadyToRun publish test for behavior already covered by… View resolved comment
Medium severity.github/​scripts/​compare_perf_timings.py — The rolling baseline still considers every prior sample with the same processor count comparable.… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — This records the runtime of MSTest.Performance.Runner, not the launched benchmark process. The… View resolved comment
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Scenarios/​Scenario1.cs — Use the requested target framework here. The generated project is patched with _tfm, but… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/MSTest.Performance.Runner/Steps/ProcessBenchmarkRunner.cs:102

  • On Unix this reads CPU metrics only after WaitForExitAsync has reaped the process, which can throw or lose the final metric. The existing ProcessMeasurement.WaitForExitAndSampleTotalProcessorTimeAsync helper explicitly samples while the process is alive for this reason (ProcessMeasurement.cs:20-33). Use that helper as the task driving the resource loop so Linux standalone lanes can reliably produce reports.
 Task exitTask = process.WaitForExitAsync();
long peakWorkingSetBytes = 0;
while (captureProcessResources && !exitTask.IsCompleted)
{
peakWorkingSetBytes = UpdatePeakWorkingSet(process, peakWorkingSetBytes);

Comment threadtest/Performance/MSTest.Performance.Runner/Steps/DotnetPublisher.cs Outdated
Quote generated project paths and reuse the cross-platform live CPU sampler so NativeAOT publishing and standalone reports remain reliable on paths with spaces and Unix hosts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 14:48

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: 1 High severity · 1 Medium severity

New issues introduced by this change (2)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the…
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw…
Issues resolved since last review (1)
SeverityFinding
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — Quote the generated project path before passing this command to DotnetCli. TargetAssetPath is… View resolved comment

Let the NativeAOT RID select the target architecture and preserve the last working-set sample when Unix process metrics disappear at exit.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:06

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Issues resolved since last review (2)
SeverityFinding
Medium severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​ProcessBenchmarkRunner.cs — The working-set sample races with process exit on Unix. PeakWorkingSet64 can throw… View resolved comment
High severitytest/​Performance/​MSTest.Performance.Runner/​Steps/​DotnetPublisher.cs — The generated scenario fixes PlatformTarget to x64 (Scenario1.cs:199), but this chooses the… View resolved comment
Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

test/Performance/README.md:44

  • The artifact cannot enforce the documented worker-count match: SingleProject/ProcessBenchmarkReport never records _workers, and load_current_results therefore cannot include it in the baseline environment key. Changing a pipeline's explicit worker count while retaining its name would silently compare incompatible runs, and users cannot verify the dimension from Result.json. Please persist the effective worker count and include it in baseline matching, or narrow this claim and require pipeline resets when worker settings change.
Compare results only when the OS and CI image, architecture, processor count,
runner runtime, target framework, configuration, worker count, and scenario all
match. The rolling baseline enforces these environment dimensions. Prefer

Persist effective MSTest worker counts in schema-v3 reports and baseline environment keys so concurrency changes cannot reuse incompatible history.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
CopilotAI review requested due to automatic review settings August 28, 2026 15:28

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

Review tier: Balanced
Findings: None

Retrigger the required pipeline after an unrelated allocation-sensitive telemetry test failed on Windows Debug.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: fd0b7d0f-8590-4c8b-ae58-635c652c60ef
@Evangelink
Amaury Levé (Evangelink) merged commit b13db97 into microsoft:mainAug 28, 2026
30 checks passed
@Evangelink
Amaury Levé (Evangelink) deleted the dev/amauryleve/compare-test-framework-performance branch August 28, 2026 19:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

state/needs-reviewAwaiting review from the team.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Evangelink@0101