Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use FillChar for debug fill patterns of 40 bytes and larger - #96

Closed
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar
Closed

Use FillChar for debug fill patterns of 40 bytes and larger#96
janrysavy wants to merge 1 commit into
pleriche:masterfrom
janrysavy:perf/debug-fillchar

Conversation

@janrysavy

@janrysavyjanrysavy commented Jul 21, 2026

Copy link
Copy Markdown

Summary

This adds a small FillChar fast path to FillFreedDebugBlockWithDebugPattern and FillAllocatedDebugBlockWithDebugPattern for user sizes of 40 bytes and larger. In the 4 KiB allocate/free benchmark, throughput improved by 52.36% to 62.89% on AMD and 104.58% to 128.59% on Intel across Win64 and Win32. Smaller blocks continue to use the original scalar code. For freed blocks, the code restores the TFastMM_FreedObject pointer after filling, preserving the existing debug-block layout and diagnostics. The diff adds 15 lines.

This is complementary to #92: that PR can disable fill patterns for profiling-oriented configurations, while this change speeds up the existing full diagnostic behavior when fill patterns remain enabled.

Tracking issue: #99

Benchmark

The benchmark repeatedly allocates and frees one block in FastMM5 debug mode with stack-trace collection disabled. Results are medians of 25 alternating baseline/candidate process pairs after three warmups. Throughput gain is (baseline time / candidate time - 1) * 100; the 95% interval is a deterministic 10,000-resample bootstrap interval for the paired median.

AMD Ryzen 9 7950X:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B4,000,000190.56 ms175.34 ms8.96% (8.46% to 10.13%)
Win3264 B4,000,000171.22 ms174.36 ms-1.62% (-3.20% to -1.07%)
Win644 KiB750,000337.03 ms221.28 ms52.36% (52.16% to 52.49%)
Win324 KiB750,000358.80 ms221.28 ms62.89% (51.07% to 65.11%)
Win6464 KiB50,000311.22 ms192.02 ms61.87% (61.72% to 62.03%)
Win3264 KiB50,000331.08 ms192.58 ms71.96% (71.59% to 73.39%)

Intel Core i7-8750H:

TargetUser sizeOperationsBaseline medianCandidate medianThroughput gain (95% CI)
Win6464 B2,100,000205.62 ms185.75 ms10.55% (9.41% to 11.92%)
Win3264 B2,000,000174.79 ms161.55 ms8.17% (7.65% to 8.94%)
Win644 KiB268,000310.53 ms151.43 ms104.58% (103.13% to 105.44%)
Win324 KiB286,000365.08 ms159.67 ms128.59% (125.41% to 129.47%)
Win6464 KiB17,000291.05 ms137.29 ms112.54% (110.16% to 114.24%)
Win3264 KiB18,000337.73 ms142.56 ms136.63% (134.74% to 137.90%)

The measured 4 KiB and 64 KiB cases improved substantially; the 64-byte controls were near the crossover, ranging from a 1.62% loss to a 10.55% gain. Additional 25-pair tests cover 1, 39, 40, and 41 bytes on both architectures.

The worst AMD result was -1.89% at one byte on Win32, with all Win64 controls positive. On Intel, the worst results were -1.62% at 39 bytes on Win32 (CI -2.37% to -0.08%) and -1.27% at 39 bytes on Win64 (CI -2.26% to -0.33%). Avoiding those small boundary losses would require duplicating or restructuring the scalar loops, which does not seem justified for debug mode.

Tests ran on an AMD Ryzen 9 7950X and an Intel Core i7-8750H under Windows 11 build 26200, using Delphi Win32/Win64 37.0.59082.6021 with release-style -O+ builds. The Intel runs use fewer operations to keep process duration similar; both builds still perform identical work within each pair, and the pairing and analysis are unchanged.

Correctness

FastMMDebugPatternTest.dpr passes for the DCC 37 baseline and candidate on Win32 and Win64 on both CPUs. Additional compatibility runs pass with DCC 35 and 36 Win32/Win64 and with DCC 37 PurePascal. The test verifies every allocated fill byte for sizes 1 through 1025, then corrupts and restores every byte for freed-block sizes 1 through 65 and 257. It accepts only the expected EInvalidPointer and requires a clean scan afterward, covering the object marker and every baseline 8/4/2/1-byte remainder path.

Full benchmark source, raw AMD and Intel samples, summaries, and correctness tests: https://github.com/janrysavy/FastMM5/tree/ed3229e678bbe4b2455b8435168442d51f430996/perf-review

Compatibility note

FastMM5 supports Delphi versions older than the DCC 35 through 37 compilers used here. FillChar is suitable for this code and does not allocate. The original scalar path remains below 40 bytes, but the crossover should still be checked with the oldest supported RTL.

@janrysavy
janrysavy marked this pull request as ready for review July 21, 2026 09:44
@pleriche

Copy link
Copy Markdown
Owner

Thanks Jan, I've implemented your suggestion. I have wrapped it in a "{$if CompilerVersion >= 28}" (XE7), because IIRC that is where FillChar saw a significant improvement. (Please let me know if I am wrong.)

@janrysavy

Copy link
Copy Markdown
Author

Pierre, could you please reference the original PRs or issues in your commits so there’s a clear link back for future reference?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@janrysavy@pleriche