Uh oh!
There was an error while loading. Please reload this page.
perf(sort): fast path for full batch when it sorts before all other streams in SortPreservingMergeStream - #23345
perf(sort): fast path for full batch when it sorts before all other streams in SortPreservingMergeStream#23345rluvaton wants to merge 11 commits into
SortPreservingMergeStream#23345Conversation
rluvaton
commented
Jul 6, 2026
run benchmarks |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (f3cffeb) to 6a0e76e (merge-base) diff using: tpcds File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (f3cffeb) to 6a0e76e (merge-base) diff using: clickbench_partitioned File an issue against this benchmark runner |
rluvaton
commented
Jul 6, 2026
run benchmark sort_preserving_merge |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (f3cffeb) to 6a0e76e (merge-base) diff using: tpch File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (f3cffeb) to 6a0e76e (merge-base) diff using: sort_preserving_merge File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagesort_preserving_merge — base (merge-base)
sort_preserving_merge — branch
File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 6, 2026
run benchmark sort_preserving_merge |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (1d5183b) to 0365d3c (merge-base) diff using: sort_preserving_merge File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagesort_preserving_merge — base (merge-base)
sort_preserving_merge — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 6, 2026
run benchmark sort_preserving_merge |
This reverts commit 5a5e1f0.
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (5a5e1f0) to 0365d3c (merge-base) diff using: sort_preserving_merge File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagesort_preserving_merge — base (merge-base)
sort_preserving_merge — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 6, 2026
run benchmark clickbench_sorted |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (80d85bf) to 0365d3c (merge-base) diff using: clickbench_sorted File an issue against this benchmark runner |
rluvaton
commented
Jul 6, 2026
run benchmark sort_tpch10 |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (80d85bf) to 0365d3c (merge-base) diff using: sort_tpch10 File an issue against this benchmark runner |
adriangbot
commented
Jul 6, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagesort_tpch10 — base (merge-base)
sort_tpch10 — branch
File an issue against this benchmark runner |
SortPreservingMergeStreamrluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
Hi @rluvaton, your benchmark configuration could not be parsed (#23345 (comment)). Error: Usage: Any benchmark name is accepted: Per-side configuration ( env:
# shared env is inherited by BOTH the build and the run, so build# flags go here. Builds default to no debuginfo for speed; opt back# in for hung-job gdb dumps and cap jobs to stay within memory:CARGO_PROFILE_RELEASE_DEBUG: "1"CARGO_BUILD_JOBS: "1"baseline:
ref: v45.0.0env:
# per-side env only reaches the benchmark run, not the buildDATAFUSION_RUNTIME_MEMORY_LIMIT: 1Gchanged:
ref: v46.0.0env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: 2GFile an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (92505cf) to 0365d3c (merge-base) diff using: tpch10 File an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
show benchmark queue |
adriangbot
commented
Jul 7, 2026
Hi @rluvaton, you asked to view the benchmark queue (#23345 (comment)).
File an issue against this benchmark runner |
adriangbot
commented
Jul 7, 2026
Benchmark for this request hit the 7200s job deadline before finishing. Benchmarks requested: Kubernetes messageFile an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (92505cf) to 0365d3c (merge-base) diff using: tpch10 File an issue against this benchmark runner |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (92505cf) to 0365d3c (merge-base) diff using: tpch10 File an issue against this benchmark runner |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (92505cf) to 0365d3c (merge-base) diff using: tpch10 File an issue against this benchmark runner |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)DetailsResource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
rluvaton
commented
Jul 7, 2026
run benchmark tpch10 |
adriangbot
commented
Jul 7, 2026
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing fast-batch-wise-skip (92505cf) to 0365d3c (merge-base) diff using: tpch10 File an issue against this benchmark runner |
adriangbot
commented
Jul 7, 2026
Benchmark for this request hit the 7200s job deadline before finishing. Benchmarks requested: Kubernetes messageFile an issue against this benchmark runner |
alamb
commented
Jul 7, 2026
Hi @rluvaton -- I am sort of lost on this PR -- it seems like it adds some non trivial complexity but I can't quite seem what benchmarks are made faster. is it really tpch_sort (aka #23345 (comment)) ? Is there any way to make this PR (and the logic) simpler? |
rluvaton
commented
Jul 8, 2026
Yes, I also created dedicated micro benchmarks:
tried to simplify it |
@alamb In order to make the code simpler I'm refactoring the code before: |
Which issue does this PR close?
N/A
Rationale for this change
this is part of an effort I'm doing to make sort much much faster.
here I'm making it faster when the data is somewhat sorted
in this there are 2 things that we save on in case we get to the fast path:
What changes are included in this PR?
fast skip when can emit entire batch
Are these changes tested?
yes and existing tests
Are there any user-facing changes?
Not API changes but behavioral changes regarding the size of the output batches:
batch_sizefor already sorted input - it will return the same size as the input streame.g. we have 3 streams and batch size is 4:
the output will be:
TODO