Uh oh!
There was an error while loading. Please reload this page.
[Optimize](Map) Optimize MapFileColumnIterator::read_by_rowids for batched map access - #58485
Merged
Merged
Conversation
hello-stephen
commented
Nov 28, 2025
Contributor
Thank you for your contribution to Apache Doris. Please clearly describe your PR:
|
eldenmoon
commented
Nov 28, 2025
MemberAuthor
run buildall |
doris-robot
commented
Nov 28, 2025
TPC-H: Total hot run time: 34379 ms |
doris-robot
commented
Nov 28, 2025
TPC-DS: Total hot run time: 184924 ms |
…tched map access - Replace the per-row seek + next_batch(1) loop in MapFileColumnIterator::read_by_rowids with batched offset reads via OffsetFileColumnIterator::read_by_rowids to derive key/value ranges for the requested rowids. - Compute per-row map sizes from offset[rowid] and offset[rowid+1], using the page-tail next_array_item_ordinal sentinel for the last row when rowid+1 is out of bounds. - Skip key/value decoding for null rows by consulting a pre-fetched null map, and add a safety check to reject non-nullable destination columns when the underlying map reader is nullable. - Reuse a small peek column in OffsetFileColumnIterator::_peek_one_offset to avoid repeated temporary column allocations when reading page sentinels. - Add a unit test (MapReadByRowidsSkipReadingResizesDestination) to verify that read_by_rowids honors the SKIP_READING flag and only resizes the destination column without touching sub-iterators.
eldenmoonforce-pushed
the
map-reader-opt
branch
from
November 28, 2025 04:40
3b9e323 to
a9b5d9dCompareeldenmoon
commented
Nov 28, 2025
MemberAuthor
run buildall |
doris-robot
commented
Nov 28, 2025
ClickBench: Total hot run time: 27.14 s |
doris-robot
commented
Nov 28, 2025
TPC-H: Total hot run time: 34259 ms |
doris-robot
commented
Nov 28, 2025
TPC-DS: Total hot run time: 182291 ms |
doris-robot
commented
Nov 28, 2025
ClickBench: Total hot run time: 27.27 s |
hello-stephen
commented
Nov 28, 2025
Contributor
BE UT Coverage ReportIncrement line coverage Increment coverage report
|
hello-stephen
commented
Nov 28, 2025
Contributor
BE Regression && UT Coverage ReportIncrement line coverage Increment coverage report
|
hello-stephen
commented
Dec 1, 2025
Contributor
BE Regression && UT Coverage ReportIncrement line coverage Increment coverage report
|
csun5285
reviewed
Dec 1, 2025
Uh oh!
There was an error while loading. Please reload this page.
Contributor
PR approved by anyone and no changes requested. |
Contributor
PR approved by at least one committer and no changes requested. |
Uh oh!
There was an error while loading. Please reload this page.
nagisa-kunhah pushed a commit
to nagisa-kunhah/doris
that referenced
this pull request
Dec 14, 2025
…tched map access (apache#58485) - Replace the per-row seek + next_batch(1) loop in MapFileColumnIterator::read_by_rowids with batched offset reads via OffsetFileColumnIterator::read_by_rowids to derive key/value ranges for the requested rowids. - Compute per-row map sizes from offset[rowid] and offset[rowid+1], using the page-tail next_array_item_ordinal sentinel for the last row when rowid+1 is out of bounds. - Skip key/value decoding for null rows by consulting a pre-fetched null map, and add a safety check to reject non-nullable destination columns when the underlying map reader is nullable. - Reuse a small peek column in OffsetFileColumnIterator::_peek_one_offset to avoid repeated temporary column allocations when reading page sentinels. - Add a unit test (MapReadByRowidsSkipReadingResizesDestination) to verify that read_by_rowids honors the SKIP_READING flag and only resizes the destination column without touching sub-iterators. - Improve performance from ~19s to ~0.1s in the worst-case access pattern, and from ~6s to ~3s in the normal case.
mrhhsg pushed a commit
that referenced
this pull request
Dec 23, 2025
…tched map access (#58485) - Replace the per-row seek + next_batch(1) loop in MapFileColumnIterator::read_by_rowids with batched offset reads via OffsetFileColumnIterator::read_by_rowids to derive key/value ranges for the requested rowids. - Compute per-row map sizes from offset[rowid] and offset[rowid+1], using the page-tail next_array_item_ordinal sentinel for the last row when rowid+1 is out of bounds. - Skip key/value decoding for null rows by consulting a pre-fetched null map, and add a safety check to reject non-nullable destination columns when the underlying map reader is nullable. - Reuse a small peek column in OffsetFileColumnIterator::_peek_one_offset to avoid repeated temporary column allocations when reading page sentinels. - Add a unit test (MapReadByRowidsSkipReadingResizesDestination) to verify that read_by_rowids honors the SKIP_READING flag and only resizes the destination column without touching sub-iterators. - Improve performance from ~19s to ~0.1s in the worst-case access pattern, and from ~6s to ~3s in the normal case.
16 tasks
yiguolei pushed a commit
that referenced
this pull request
Dec 24, 2025
…rning (#59286) ### What problem does this PR solve? Problem Summary: ### Release note Cherry-pick #58370#58354#59043#58851#58485#58682#58614#58373#57204#58719#58471#58573#58657 ### Check List (For Author) - Test <!-- At least one of them must be included. --> - [ ] Regression test - [ ] Unit Test - [ ] Manual test (add detailed scripts or steps below) - [ ] No need to test or manual test. Explain why: - [ ] This is a refactor/code format and no logic has been changed. - [ ] Previous test can cover this change. - [ ] No code files have been changed. - [ ] Other reason <!-- Add your reason? --> - Behavior changed: - [ ] No. - [ ] Yes. <!-- Explain the behavior change --> - Does this need documentation? - [ ] No. - [ ] Yes. <!-- Add document PR link here. eg: apache/doris-website#1214 --> ### Check List (For Reviewer who merge this PR) - [ ] Confirm the release note - [ ] Confirm test cases - [ ] Confirm document - [ ] Add branch pick label <!-- Add branch pick label that this PR should merge into --> --------- Co-authored-by: 924060929 <lanhuajian@selectdb.com> Co-authored-by: Jerry Hu <mrhhsg@gmail.com> Co-authored-by: Jerry Hu <hushenggang@selectdb.com> Co-authored-by: lihangyu <lihangyu@selectdb.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
from ~6s to ~3s in the normal case.
What problem does this PR solve?
Issue Number: close #xxx
Related PR: #xxx
Problem Summary:
Release note
None
Check List (For Author)
Test
Behavior changed:
Does this need documentation?
Check List (For Reviewer who merge this PR)