Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat: optimize XNode memory layout and refactor XEntry type resolution by fslongjin · Pull Request #14 · asterinas/xarray · GitHub
Skip to content

feat: optimize XNode memory layout and refactor XEntry type resolution - #14

Open
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor
Open

feat: optimize XNode memory layout and refactor XEntry type resolution#14
fslongjin wants to merge 1 commit into
asterinas:mainfrom
fslongjin:feature/slab-friendly-and-entry-refactor

Conversation

@fslongjin

@fslongjinfslongjin commented Feb 28, 2026

Copy link
Copy Markdown

Summary

  • Slab-friendly XNode memory layout (optional): Adds a slab-friendly feature. When enabled, XNode's slots field is heap-allocated via Box<[XEntry; 64]> instead of inline. This changes the allocation footprint seen by slab-style allocators: the original ~544B node may round up to a 1K size class; with the feature, the node header and the slot array can fit smaller classes (e.g. ~64B + 512B), potentially reducing internal fragmentation in slab-heavy environments. Note: Under a default/general-purpose allocator, benchmarks in this repo do not show a net speedup; the feature is an allocator-dependent trade-off and can be beneficial in some setups (e.g. custom slab allocators).
  • Pin smallvec to 1.13.1 for reproducible builds.
  • Refactor XEntry::ty(): Replace chained not()/then() with a clear if/else for readability.
  • Debuggability: Add a descriptive panic message for invalid XEntry tags.
  • Cleanup: Remove unused Not import from core::ops.
  • Benchmarks: Add benches/xarray_bench.rs for cargo bench (store/load cursor, random load, range iteration, COW clone+overwrite). Run with --features std and optionally --features "std slab-friendly" to compare.

Usage

Enable the feature in Cargo.toml if you want the slab-friendly layout:

[dependencies]
xarray = { version = "...", features = ["slab-friendly"] }

Run benchmarks:

cargo bench --features std --bench xarray_bench
cargo bench --features "std slab-friendly" --bench xarray_bench

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 7f5589e to 372a5ffCompareFebruary 28, 2026 07:46
@tatetian

tatetian commented Feb 28, 2026

Copy link
Copy Markdown

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

- Add optional `slab-friendly` feature: when enabled, XNode slots are
heap-allocated via Box, changing allocation footprint for slab
allocators (may reduce internal fragmentation; allocator-dependent).
- Pin smallvec to 1.13.1 for reproducibility.
- Simplify XEntry::ty() with explicit if-else; add panic message for
invalid XEntry tags; remove unused Not import.
- Add benches/xarray_bench.rs for cargo bench (store, cursor load,
random load, range, COW clone+overwrite).
Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Author

Thanks for trying to contribute.

  • reducing the effective allocation size from the previous 544-byte (which rounds up to 1K in slab allocators) to a slab-friendly 576 bytes.

Reduce from 544B to 576B?

  • This also improves cache performance due to better cacheline alignment.

How exactly is the cache performance improved?

This optimization introduced this PR has no empirical result to back up. So I am not convinced that this slab-friendly feature leads to a net positive result. And if this change is a net positive, why gated by a feature? It should simply replace the current implementation.

In addition, from the PR description, it seems that this PR should be broken into several atomic commits.

Lastly, I am not sure that we want to main this crate anymore. This crate is not updated for quite some time. The latest implementation of Xarray has been moved into the Asterinas mono repo for better integration with RCU.

I apologize — the previous commit message incorrectly described the performance impact of the slab-friendly change. It should not have claimed that the optimization improves benchmark performance in general; the trade-off is allocator- and workload-dependent. I have amended the commit and updated the PR description accordingly.

Thanks for the careful review. Let me address the two concerns separately.

1) "Reduce from 544B to 576B?"

You are right that the wording is confusing if interpreted as raw struct size only.

What I intended is effective allocation footprint under slab allocators, not only size_of::<XNode>():

  • In the original layout, XNode contains inline slots, so one node object is ~544B.
  • In many slab allocators with coarse size classes (e.g. power-of-two), a 544B object may be placed in a 1KiB class.
  • With slab-friendly, we split node metadata and slots:
    • node header (small object, typically one small slab class),
    • slot array (Box<[XEntry; 64]>, 512B class).
  • So the allocator-visible total can become closer to small-class + 512B (for example ~64B + 512B = ~576B), which reduces internal fragmentation versus a single 1KiB class allocation.

So the key claim is about allocator class fit and internal fragmentation, not that a single Rust object shrinks from 544 to 576.

2) "Where is the empirical evidence?"

I ran cargo bench in this branch with the same benchmark suite under two feature sets:

  • --features std
  • --features "std slab-friendly"

Environment: local Linux machine(ubuntu 24.04, i5-12450H), default allocator, same code, same benchmark target (benches/xarray_bench.rs).

Benchmarkstd (ns/iter)std+slab-friendly (ns/iter)Delta
bench_store_dense8,517,270.358,717,937.25+2.36%
bench_cursor_load_dense603,782.84638,052.65+5.68%
bench_load_random_dense3,843,162.604,415,553.70+14.89%
bench_range_sparse_even1,420,373.751,489,634.25+4.88%
bench_cow_clone_then_overwrite4,475,489.254,472,778.65-0.06%

I apologize — the earlier “slight performance improvement” claim was based on measurements under DragonOS’s memory allocator. Thank you for pointing out the lack of empirical evidence. I have added the benchmark suite above; on Rust’s default allocator the numbers do show a performance regression for slab-friendly in these tests.

I have also removed that claim from the PR description and commit message.

@fslongjin
fslongjinforce-pushed the feature/slab-friendly-and-entry-refactor branch from 372a5ff to ada14e6CompareFebruary 28, 2026 17:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@fslongjin@tatetian