Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); std.sort.equalRange: improve performance by using center-left upperBound by Olvilock · Pull Request #21278 · ziglang/zig · GitHub
Skip to content

std.sort.equalRange: improve performance by using center-left upperBound - #21278

Closed
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master
Closed

std.sort.equalRange: improve performance by using center-left upperBound#21278
Olvilock wants to merge 0 commit into
ziglang:masterfrom
Olvilock:master

Conversation

@Olvilock

Copy link
Copy Markdown

Ports solution presented at CppCon by Andrei Alexandrescu.

The upper bound uses result of the lower bound, and also skews the first few choppings to the left, leading up to 1.5x speedup (as reported by Andrei)

@LucasSantos91

Copy link
Copy Markdown
Contributor

Andrei's algorithm assumes that the result will most frequently be found in the lower indices. I watched the talk a long time ago, and, from what I remember, Andrei claimed that this is the most common case, but he never backed it up with evidence. In my project, I use binary search in two places and, in both cases, that assumption does not hold, and this change would therefore decrease performance. I think it would be better to leave binary search as is, and create a new function, called, for instance, leftBiasedBinarySearch, that implements this algorithm.

@rohlem

rohlem commented Sep 2, 2024

Copy link
Copy Markdown
Contributor

@LucasSantos91 From how I understood the talk, the assumption for the portion optimizing general binary search is that the input value can lie outside of the range of contained elements (also referred to as an "open range search" iiuc, otherwise it would be referred to as a "closed range search").
The resulting algorithm was to compare the center element first, which gives information on whether the element to search for is in the lower or higher half, then skew the binary search towards the end/extreme until a more extreme element is found to close the search range, after which it falls back to classic unskewed binary search, since this is more optimal for the remaining closed range search.
I currently can't connect your interpretation / statement about lower indices with anything I remember from the talk, feel free to clarify (maybe with a timestamp).

Regardless, this PR only seems to port the equalRange improvement, which is compared under the same closed-range assumption that is already the case with the previous implementation. The normal binarySearch function is unmodified.
The algorithm starts with normal, non-skewed binarySearch, and only skews the relative upper bound to the (center) left, because it's generally (statistically) less likely for the same element to repeat for a majority of the rest of the sequence. Once an element inside the range is found, it reverts to non-skewed binarySearch for the remaining upper bound within the closed range.

Some sort of performance benchmarking code and results as motivation for whether to merge this would be nice.

@LucasSantos91

Copy link
Copy Markdown
Contributor

Sorry, my bad. I was thinking of binary search instead of equal range.
For the equal range, I think the biggest performance gain comes from doing the lower bound simultaneously with the upper bound. This way, you only need to iterate over the range once, as each step gives you information that can be used to tighten one end of the resulting range.

@Olvilock

Olvilock commented Sep 3, 2024

Copy link
Copy Markdown
Author

I've written a little benchmark to assess which approach would be best in the case of std.sort.equalRange. I compare the current implementation with this proposal and 21290. To reproduce the results, curl this words.txt file.
My results (using the latest zig master):
$ zig build-exe bench.zig -O ReleaseSafe && ./bench
Vanilla function: 562 ms
Center-left function: 374 ms
Simul function: 271 ms
Totals: { 4665490000, 0, 4665490000 }

$ zig build-exe bench.zig -O ReleaseFast && ./bench
Vanilla function: 380 ms
Center-left function: 240 ms
Simul function: 252 ms
Totals: { 4665490000, 0, 4665490000 }

It seems that there's a bug somewhere in my version, and it is also slower than 21290, so this pull definitely needs holding off. Not closing yet because 21290 needs two comparisons to work during the common iterations, and therefore can be slower than lowerBound + adjusted upperBound for trivial types.

@OlvilockOlvilock closed this Sep 5, 2024
@Olvilock

Copy link
Copy Markdown
Author

It seems to me that under the current API, this implementation is bound to underperform in ReleaseSafe, because the unreachable is not able to optimize out the redundant checks. Therefore I do not see this as a better implementation unless the API gets reverted to 0.13. I will reimplement this pull under 0.13 API and reopen later

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Olvilock@LucasSantos91@rohlem