ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x) - #8191

Closed
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster
Closed

ARROW-10010: [Rust] Speedup arithmetic (1.3-1.9x)#8191
jorgecarleitao wants to merge 4 commits into
apache:masterfrom
jorgecarleitao:divide_simd_faster

Conversation

@jorgecarleitao

Copy link
Copy Markdown
Member

This PR speeds-up arithmetic ops by leveraging vectorization of non-divide operations (in non-SIMD), as well as removing an un-needed operation in SIMD division.

For non-SIMD, this yields about [-30%,-45%] for all operations (+-*/)
For SIMD, this yields about -30% on division.

The culprit in non-SIMD was that we required the operation to return Result<T::Native>, which was not allowing the compiler to vectorize the operation. Only the division requires Result. For divide, removing the operator further speed up the operation (I do not know the reason).

The culprit in SIMD was primarily a simd_load too many that was not doing anything.

Benchmarks

The benchmark used:

set -e
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels
echo "##################################"
git checkout 0852869d1a9b7da4a1b91fa7cb7d4ef48e99cdec
cargo bench --bench arithmetic_kernels --features simd
git checkout divide_simd_faster
cargo bench --bench arithmetic_kernels --features simd

and below are the results for the execution of the second bench, which is the one that gives the differential, in my machine:

Non-SIMD

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 37.24s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-d281862a43faaf38
Gnuplot not found, using plotters backend
add 512 time: [1.4714 us 1.4758 us 1.4803 us] change: [-44.446% -43.969% -43.522%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
subtract 512 time: [1.4825 us 1.4844 us 1.4866 us] change: [-45.351% -45.018% -44.686%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
5 (5.00%) high mild
4 (4.00%) high severe
multiply 512 time: [1.4895 us 1.4936 us 1.4990 us] change: [-44.822% -44.135% -43.479%] (p = 0.00 < 0.05)
Performance has improved.
Found 9 outliers among 100 measurements (9.00%)
4 (4.00%) high mild
5 (5.00%) high severe
divide 512 time: [1.9742 us 1.9773 us 1.9810 us] change: [-33.273% -32.688% -32.052%] (p = 0.00 < 0.05)
Performance has improved.
Found 14 outliers among 100 measurements (14.00%)
7 (7.00%) high mild
7 (7.00%) high severe
limit 512, 512 time: [374.66 ns 375.64 ns 376.53 ns] change: [-0.1000% +0.4442% +0.9503%] (p = 0.10 > 0.05)
No change in performance detected.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) low severe
2 (2.00%) low mild
2 (2.00%) high mild
2 (2.00%) high severe
add_nulls_512 time: [1.4880 us 1.4982 us 1.5115 us] change: [-44.084% -43.116% -42.111%] (p = 0.00 < 0.05)
Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
3 (3.00%) high mild
13 (13.00%) high severe
divide_nulls_512 time: [1.9731 us 1.9758 us 1.9790 us] change: [-33.404% -32.570% -31.416%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

SIMD

divide is the only relevant

Previous HEAD position was 0852869d1 Improved benches for arithmetic.
Switched to branch 'divide_simd_faster'
Compiling arrow v2.0.0-SNAPSHOT (/Users/jorgecarleitao/projects/arrow/rust/arrow)
Finished bench [optimized] target(s) in 38.63s
Running /Users/jorgecarleitao/projects/arrow/rust/target/release/deps/arithmetic_kernels-b8dc1739cfb5ae36
Gnuplot not found, using plotters backend
add 512 time: [879.31 ns 883.95 ns 889.17 ns] change: [-0.2041% +0.6502% +1.5484%] (p = 0.15 > 0.05)
No change in performance detected.
Found 16 outliers among 100 measurements (16.00%)
5 (5.00%) high mild
11 (11.00%) high severe
subtract 512 time: [864.99 ns 866.95 ns 868.95 ns] change: [-4.8531% -4.1561% -3.5163%] (p = 0.00 < 0.05)
Performance has improved.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe
multiply 512 time: [862.85 ns 864.87 ns 867.71 ns] change: [-3.8532% -3.1774% -2.4459%] (p = 0.00 < 0.05)
Performance has improved.
Found 5 outliers among 100 measurements (5.00%)
5 (5.00%) high severe
divide 512 time: [1.9703 us 1.9771 us 1.9843 us] change: [-30.046% -29.457% -28.903%] (p = 0.00 < 0.05)
Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
1 (1.00%) high severe
limit 512, 512 time: [368.89 ns 369.96 ns 370.96 ns] change: [-1.9574% -1.0063% -0.0347%] (p = 0.04 < 0.05)
Change within noise threshold.
Found 26 outliers among 100 measurements (26.00%)
5 (5.00%) low severe
6 (6.00%) low mild
9 (9.00%) high mild
6 (6.00%) high severe
add_nulls_512 time: [871.97 ns 876.99 ns 883.57 ns] change: [-5.1106% -3.6889% -2.3080%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe
divide_nulls_512 time: [1.9582 us 1.9625 us 1.9678 us] change: [-34.188% -33.161% -32.136%] (p = 0.00 < 0.05)
Performance has improved.
Found 8 outliers among 100 measurements (8.00%)
2 (2.00%) high mild
6 (6.00%) high severe

@github-actions

Copy link
Copy Markdown

let null_bit_buffer =
combine_option_bitmap(left.data_ref(), right.data_ref(), left.len())?;
let bitmap = null_bit_buffer.map(Bitmap::from);
let bitmap = null_bit_buffer.clone().map(Bitmap::from);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this clone necessary?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe not, but I was unable to get a bitmap reference to set the mask for SIMD from a buffer without clone.

@jorgecarleitao
jorgecarleitao deleted the divide_simd_faster branch September 15, 2020 09:11
@asfimportasfimport mentioned this pull request Sep 15, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jorgecarleitao@nevi-me