GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice - #46730

Closed
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677
Closed

GH-46677: [C++] Expose an BinaryViewBuilder interface for append a binary and multiple subslice#46730
IndifferentArea wants to merge 12 commits into
apache:mainfrom
IndifferentArea:GH-46677

Conversation

@IndifferentArea

@IndifferentAreaIndifferentArea commented Jun 6, 2025

Copy link
Copy Markdown
Contributor

Rationale for this change

see #46677

What changes are included in this PR?

see #46677

Are these changes tested?

Yes

Are there any user-facing changes?

No

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

@mapleFU is currently implemented interface expected?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBlock(value.data(), static_cast<int64_t>(value.size()));
}

Status AppendViewFromBuffer(int32_t buffer_id, int32_t buffer_offset, int32_t start,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

naming: from buffer or from block?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Personally both is ok for me, I prefer Buffer since a variable is buffer_index

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
UnsafeAppend(value.data(), static_cast<int64_t>(value.size()));
}

Result<std::pair<int32_t, int32_t>> AppendBlock(const uint8_t* value,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use more specific name rather than pair<i32, i32>?

@IndifferentAreaIndifferentAreaJun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can i directly use BinaryViewType::c_type since it already contains these two info we need?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The syntax is a bit weird here? Append a BinaryView and then append the sub-slice of the view?


Result<std::pair<int32_t, int32_t>> BinaryViewBuilder::AppendBlock(const uint8_t* value,
const int64_t length) {
DCHECK_GT(length, TypeClass::kInlineSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If length <= kInlineSize, should this return false or ok? Why just DCHECK here?

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
c_type GetViewFromBlock(int32_t block_id, int32_t block_offset, int32_t offset,
int32_t length) const {
const auto* value = blocks_.at(block_id)->data_as<uint8_t>() + block_offset + offset;
if (length <= BinaryViewType::kInlineSize) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uses ToBinaryView?

@github-actionsgithub-actionsBot added awaiting committer review Awaiting committer review and removed awaiting review Awaiting review labels Jun 7, 2025
@IndifferentArea

IndifferentArea commented Jun 7, 2025

Copy link
Copy Markdown
ContributorAuthor

Should we rename AppendBuffer or redesign the interface's semantics? since currently we don't really append a buffer/block, we will directly append to the last one if remaining size is enough.. I think It may introduce confusion.

Maybe aligning with arrow-rs's impl is fine..

@mapleFU

Copy link
Copy Markdown
Member

Some personal thoughts:

  1. AppendBuffer ( and etc ) which returns a StringView is a bit weird, Block is not a StringView
  2. Now aligned with arrow-rs also a good way

@IndifferentArea

Copy link
Copy Markdown
ContributorAuthor

Not sure why these 2 ci always failed..

@IndifferentArea
IndifferentArea marked this pull request as ready for review June 8, 2025 12:44
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
Comment threadcpp/src/arrow/array/array_test.cc Outdated
@IndifferentArea

IndifferentArea commented Jun 14, 2025

Copy link
Copy Markdown
ContributorAuthor

To implement interface aligned with what arrow-rs did, i have to change some behavior of StringHeapBuild::FinishLastBlock() and mark it as public.

More specific, before FinishLastBlock() just resize the last block. Now FinishLastBlock() reset internal states, including current_offset_, current_out_buffer_ and current_remaining_bytes_. I believe it's more safe and reasonable.
The interface was private before so don't mind external usage, for current internal usage of FinishLastBlock(), they always reset or change all these status.

If this change is unacceptable, plz let me know, i'll try to find another way.

@mapleFUmapleFU left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

General LGTM. Also cc @pitrou for your advices for this interface. This interface would be used for read binary as stringView from bytearray type faster

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
return AppendBuffer(reinterpret_cast<const uint8_t*>(value), length);
}

Result<int32_t> AppendBuffer(const std::string& value) {

@mapleFUmapleFUJun 14, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can std::string_view being used rather than const std::string&?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we remove this one since std::string_view is added here?

Comment threadcpp/src/arrow/array/builder_binary.h
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
Comment threadcpp/src/arrow/array/builder_binary.h Outdated
void UnsafeAppendViewFromBuffer(const int32_t buffer_idx, const int32_t start,
const int32_t length) {
UnsafeAppendToBitmap(true);
const auto v = data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
const auto v=data_heap_builder_.GetViewFromBuffer<false>(buffer_idx, start, length);
const auto v=data_heap_builder_.GetViewFromBuffer</*Safe=*/false>(buffer_idx, start, length);

Comment threadcpp/src/arrow/array/builder_binary.h Outdated
const int32_t length) {
ARROW_RETURN_NOT_OK(Reserve(1));
UnsafeAppendToBitmap(true);
ARROW_ASSIGN_OR_RAISE(const auto v, data_heap_builder_.GetViewFromBuffer<true>(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer<true>(
ARROW_ASSIGN_OR_RAISE(constautov, data_heap_builder_.GetViewFromBuffer</*Safe=*/true>(

@mapleFU
mapleFU requested a review from pitrouJune 20, 2025 15:38
@mapleFU

Copy link
Copy Markdown
Member

Gentle ping @pitrou

@pitrou

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@mapleFU

Copy link
Copy Markdown
Member

Isn't this approach wasteful? If you have lots of strings <= 12 bytes, you will still store their contents in a data buffer, while they're inlined in the string views.

@pitrou I suppose this is used to append a whole parquet page and add buffer for it

@pitrou

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

@mapleFU

Copy link
Copy Markdown
Member

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

Maybe we can also try to pick this way when average len is huge enough?

@pitrou

Copy link
Copy Markdown
Member

A nice question. I agree when most Parquet strings are <= 12 bytes, it would be memory wasted because a huge memcpy is applied. But when read large binary it would benefit a lot from this. I think usally a large memcpy might much faster than little un-continogous memcpy

That's true, but another cost is to create the views themselves. It would be nice if a prototype could tell us which speedup we can expect.

Maybe we can also try to pick this way when average len is huge enough?

Yes, that's definitely a possibility.

@mapleFU

Copy link
Copy Markdown
Member

So can we start to review this? We can set a ratio when average length > 12 or > 20

@andishgar

Copy link
Copy Markdown
Contributor

@mapleFU@pitrou
I believe this pull request is related to several other PRs I've submitted. Here's a summary:

1- API and Handling of the Last Buffer
In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Overall, my suggestion is to either modify this pull request or create a new API to support buffer sharing. It is possible to decide whether a created array should be compacted based on some metadata, in order to avoid memory bloat.

@mapleFU

Copy link
Copy Markdown
Member

Looks (1) would work in buffer style api, but for parquet reader, it might append buffer one by one.

The (2) is a good way for compute, but here I don't know the best way to handle this: whether to adaptive read it, or just throw it to "cast" or "compact". I prefer handling this in reader, and the later handling can "compact" the data when output or throwing to compute

@IndifferentArea

IndifferentArea commented Jun 30, 2025

Copy link
Copy Markdown
ContributorAuthor

Besides parquet related issue, there is NO buffer sharing mechanism in current api.

  • for large string sharing, sharing through this interface will save a lot of memory usage
  • for inlined string it costs as much as directly Append.

These new interfaces won't introduce more cost for current interfaces from my view.
I understand tradeoff on parquet as discussed above, but I believe this kind of buffer sharing interface is missing in current implementation, and there is definitely more need for this api apart from parquet.

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

@pitrou

Copy link
Copy Markdown
Member

Maybe we can open another issue/PR to discuss specifically whether/how parquet should use this api on appending a huge page and append view from it?

That sounds fair to me.

@pitrou

Copy link
Copy Markdown
Member

1- API and Handling of the Last Buffer In this pull request, I demonstrated that it’s possible to share buffers without copying or finalizing the last buffer. This avoids relocating the buffer to remove blank space, which can be a costly operation when the unused space exceeds 64 bytes.

2-

Is it a win, though? If most Parquet strings are <= 12 bytes we would pointlessly waste space and CPU time.

In this pull request, I proposed a method that could help avoid memory bloat when buffers are shared. Additionally, in this issue, I think this metadata could help determine when CompactArray should be called.

Thanks for the reminder, and sorry that this is taking a long time :) I propose that we review these PRs one by one. I've started with the CompactArray one and, once that is done, I would like to then move to the AppendArraySlice improvement.

This PR here is slightly more contentious so I think we should tackle it only after the other APIs have settled semantics.

@github-actions

Copy link
Copy Markdown

Thank you for your contribution. Unfortunately, this pull request has been marked as stale because it has had no activity in the past 365 days. Please remove the stale label or comment below, or this PR will be closed in 14 days. Feel free to re-open this if it has been closed in error. If you do not have repository permissions to reopen the PR, please tag a maintainer.

@github-actionsgithub-actionsBot added the Status: stale-warning Issues and PRs flagged as stale which are due to be closed if no indication otherwise label Jul 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting committer reviewAwaiting committer reviewComponent: C++Status: stale-warningIssues and PRs flagged as stale which are due to be closed if no indication otherwise

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@IndifferentArea@mapleFU@pitrou@andishgar