InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

InternPool: begin conversion to thread-safe data structure - #20528

Merged
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip
Jul 8, 2024
Merged

InternPool: begin conversion to thread-safe data structure#20528
andrewrk merged 15 commits into
ziglang:masterfrom
jacobly0:tsip

Conversation

@jacobly0

Copy link
Copy Markdown
Member

This is another step towards first putting codegen on a separate thread from sema and later making both sema and codegen multi-threaded. Currently, the multi-threaded features of the new data structure are disabled by InternPool.want_multi_threaded because they have a large performance impact on the compiler and are not yet in use. It is still undecided how to minimize the impact of

The main conflicting change is replacing *Zcu with Zcu.PerThread in any code that wants to be able to mutate the intern pool. I chose to pass this information around everywhere instead of attempting something with threadlocal variables because in the fast case that would turn the intern pool into a global singleton, and in the slow case would introduce an extra lookup on every intern pool mutation. Interestingly, if there is a desire to remove intern pool mutation from the backends, the per thread change can be reverted over time in parts of the code and that guarantees that the mutations can no longer in that section.

The current performance impact is in the range where I expect to be able to recoup the loss with more work, because while timing individual commits, this is close to the same time as before the change which ended up having the worst performance impact, which I mitigated by disabling features of the data structure as mentioned above. Even with this current slowdown, there may be a desire to get some or all of this change merged sooner to reduce future conflicts, and continue work separately.

Benchmark 1 (28 runs): master/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.43s ± 69.4ms 4.28s … 4.52s 1 ( 4%) 0%
peak_rss 323MB ± 1.06MB 320MB … 325MB 1 ( 4%) 0%
cpu_cycles 23.1G ± 233M 22.8G … 23.5G 0 ( 0%) 0%
instructions 45.0G ± 34.3K 45.0G … 45.0G 2 ( 7%) 0%
cache_references 2.13G ± 18.6M 2.11G … 2.19G 2 ( 7%) 0%
cache_misses 86.9M ± 1.15M 84.5M … 89.7M 2 ( 7%) 0%
branch_misses 85.7M ± 577K 84.6M … 87.1M 2 ( 7%) 0%
Benchmark 2 (26 runs): tsip/bin/zig build-exe -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=src/main.zig -Maro=lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=lib/compiler/aro_translate_c.zig -Mbuild_options=options.zig -fno-emit-bin
measurement mean ± σ min … max outliers delta
wall_time 4.68s ± 54.9ms 4.50s … 4.74s 5 (19%) 💩+ 5.7% ± 0.8%
peak_rss 324MB ± 682KB 323MB … 325MB 0 ( 0%) + 0.4% ± 0.2%
cpu_cycles 23.3G ± 168M 23.1G … 23.6G 0 ( 0%) + 0.9% ± 0.5%
instructions 40.1G ± 32.4K 40.1G … 40.1G 1 ( 4%) ⚡- 11.0% ± 0.0%
cache_references 2.15G ± 15.4M 2.13G … 2.20G 2 ( 8%) + 1.2% ± 0.4%
cache_misses 79.2M ± 1.55M 77.0M … 82.8M 0 ( 0%) ⚡- 8.9% ± 0.9%
branch_misses 73.4M ± 564K 72.6M … 74.8M 0 ( 0%) ⚡- 14.3% ± 0.4%

@andrewrk
andrewrk merged commit ab4eeb7 into ziglang:masterJul 8, 2024
@andrewrkandrewrk added the release notes This PR should be mentioned in the release notes. label Jul 8, 2024
@andrewrk

Copy link
Copy Markdown
Member

Data point: building the zig compiler with the x86_64 backend with codegen threading enabled:

--- a/src/InternPool.zig+++ b/src/InternPool.zig@@ -93,7 +93,7 @@ files: std.AutoArrayHashMapUnmanaged(Cache.BinDigest, OptionalDeclIndex) = .{},
/// Whether a multi-threaded intern pool is useful.
/// Currently `false` until the intern pool is actually accessed
/// from multiple threads to reduce the cost of this data structure.
-const want_multi_threaded = false;+const want_multi_threaded = true;
/// Whether a single-threaded intern pool impl is in use.
pub const single_threaded = builtin.single_threaded or !want_multi_threaded;
diff --git a/src/target.zig b/src/target.zig
index 2accc100b8..e236cea616 100644
--- a/src/target.zig+++ b/src/target.zig@@ -572,6 +572,7 @@ pub inline fn backendSupportsFeature(backend: std.builtin.CompilerBackend, compt
else => false,
},
.separate_thread => switch (backend) {
+ .stage2_x86_64 => true,
else => false,
},
};

before: 8f20e81
after: ab4eeb7
(before/after this PR merged)

Benchmark 1 (3 runs): /home/andy/dev/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/dev/zig/src/main.zig -Maro=/home/andy/dev/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/dev/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/dev/zig/.zig-cache/c/ecf36f9c39c1cbfb3043b07b8688debd/options.zig --cache-dir /home/andy/dev/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 12.8s ± 189ms 12.7s … 13.1s 0 ( 0%) 0%
peak_rss 403MB ± 463KB 403MB … 404MB 0 ( 0%) 0%
cpu_cycles 64.3G ± 124M 64.2G … 64.5G 0 ( 0%) 0%
instructions 133G ± 171K 133G … 133G 0 ( 0%) 0%
cache_references 4.07G ± 4.88M 4.06G … 4.07G 0 ( 0%) 0%
cache_misses 419M ± 3.33M 416M … 423M 0 ( 0%) 0%
branch_misses 435M ± 831K 434M … 436M 0 ( 0%) 0%
Benchmark 2 (3 runs): /home/andy/src/zig/build-release/stage3/bin/zig build-exe -fno-llvm -fno-lld --stack 33554432 -fno-sanitize-thread -ODebug --dep aro --dep aro_translate_c --dep build_options -Mroot=/home/andy/src/zig/src/main.zig -Maro=/home/andy/src/zig/lib/compiler/aro/aro.zig --dep aro -Maro_translate_c=/home/andy/src/zig/lib/compiler/aro_translate_c.zig -Mbuild_options=/home/andy/src/zig/.zig-cache/c/f18315ac6d459d8944cdba2c81e890df/options.zig --cache-dir /home/andy/src/zig/.zig-cache --global-cache-dir /home/andy/.cache/zig --name zig
measurement mean ± σ min … max outliers delta
wall_time 8.57s ± 66.9ms 8.52s … 8.64s 0 ( 0%) ⚡- 33.3% ± 2.5%
peak_rss 517MB ± 3.73MB 513MB … 520MB 0 ( 0%) 💩+ 28.1% ± 1.5%
cpu_cycles 65.3G ± 131M 65.2G … 65.4G 0 ( 0%) 💩+ 1.5% ± 0.4%
instructions 138G ± 1.22M 138G … 138G 0 ( 0%) 💩+ 3.5% ± 0.0%
cache_references 4.01G ± 17.6M 4.00G … 4.03G 0 ( 0%) - 1.4% ± 0.7%
cache_misses 220M ± 4.17M 215M … 223M 0 ( 0%) ⚡- 47.5% ± 2.0%
branch_misses 340M ± 2.84M 338M … 343M 0 ( 0%) ⚡- 21.9% ± 1.1%

@jacobly0
jacobly0 deleted the tsip branch July 8, 2024 21:36
@Jarred-Sumner

Jarred-Sumner commented Jul 9, 2024

Copy link
Copy Markdown
Contributor

wow 12.8s end-to-end before

with -fno-emit-bin, bun currently takes ~11s (it's about 1m 15s end-to-end, most of which is the "Emit LLVM" step)

time.cache/zig/zigbuildcheck--summarynew--verboseinfo: zigcompilerv0.13.0/root/bun/.cache/zig/zigbuild-obj-freference-trace=16-fno-strip-fno-stack-check-fno-stack-protector-fno-omit-frame-pointer-fvalgrind-fPIC-ODebug-targetnative-native-gnu.2.27-mcpunative--depasync_io--depzlib-internal--depasync--depZigGeneratedClasses--depResolvedSourceTag--depbuild_options-Mroot=/root/bun/root.zig-Masync_io=/root/bun/src/io/io_linux.zig-Mzlib-internal=/root/bun/src/deps/zlib.posix.zig-Masync=/root/bun/src/async/posix_event_loop.zig-MZigGeneratedClasses=/root/bun/build/codegen/ZigGeneratedClasses.zig-MResolvedSourceTag=/root/bun/build/codegen/ResolvedSourceTag.zig-Mbuild_options=/root/bun/.zig-cache/c/e16e8fad5a41dcf3d3a150b03d8ff201/options.zig-lc++-lc-fno-emit-bin-fformatted-panics--eh-frame-hdr--emit-relocs-ffunction-sections--cache-dir/root/bun/.zig-cache--global-cache-dir/root/.cache/zig--namebun-debug-fno-compiler-rt--zig-lib-dir/root/bun/src/deps/zig/lib--listen=-BuildSummary: 3/3stepssucceededchecksuccess└─zigbuild-objbun-debugDebugnative-native-gnu.2.27success11sMaxRSS:563M└─optionssuccess________________________________________________________Executedin14.32secsfishexternalusrtime14.19secs0.00micros14.19secssystime0.74secs482.00micros0.74secs

any ideas immediately come to mind why bun's compilation time is much slower than the zig compiler despite less code?

image

@andrewrk

andrewrk commented Jul 9, 2024

Copy link
Copy Markdown
Member

Note that my 12.8s data point above is with the x86_64 backend. You can try enabling it, but it's not the default yet due to machine code quality, debug info correctness, and behavior test coverage.

zig/build.zig

Lines 218 to 220 in 854e86c

constuse_llvm=b.option(bool, "use-llvm", "Use the llvm backend");
exe.use_llvm=use_llvm;
exe.use_lld=use_llvm;

If you're on new Apple hardware you'll have to wait for the aarch64 backend to get a similar speedup.

@mluggmlugg mentioned this pull request Jul 15, 2024
40 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release notesThis PR should be mentioned in the release notes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jacobly0@andrewrk@Jarred-Sumner