ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters - #13589

Merged
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004
Jul 26, 2022
Merged

ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters #13589
lidavidm merged 8 commits into
apache:masterfrom
lidavidm:arrow-17004

Conversation

@lidavidm

Copy link
Copy Markdown
Member

This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC.

@lidavidm

lidavidm commented Jul 12, 2022

Copy link
Copy Markdown
MemberAuthor

TODOs:

  • Add documentation
  • Are we handling time/timestamp types properly when time zones come into play?
  • Add date support as well

@lidavidm
lidavidm marked this pull request as ready for review July 13, 2022 16:47
@lidavidm

Copy link
Copy Markdown
MemberAuthor

CC @toddfarmer@lwhite1

Not all types are supported here but a core set is. If the approach looks reasonable I can extend coverage to at least Decimal and Binary types.

@lwhite1

Copy link
Copy Markdown
Contributor

Overall, this looks really nice to me.

One minor nit (not from the current change set): The statement on line 87 "Currently, it is not possible to define a custom type conversion for a supported or unsupported type." has me scratching my head a bit Should it just say "it is not possible to define a custom type conversion"? If there's a quick re-phrasing that helps, it might be worth adding.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Thanks for taking a look!

Updated the docs, and implemented binders for binary types and decimals.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

@emkornfield@liyafan82@pitrou any opinions on having this functionality (binding Arrow data -> JDBC prepared statement parameters) here?

@pitrou

Copy link
Copy Markdown
Member

Hmm, I'm out of my depth here, but I guess it looks useful? The main downside being the one-row-at-a-time mechanics, I suppose.

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm
I find an example for executeUpdate, does it also support executeQuery?

@lidavidm

Copy link
Copy Markdown
MemberAuthor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

@liyafan82

Copy link
Copy Markdown
Contributor

Interesting work! Thanks. @lidavidm I find an example for executeUpdate, does it also support executeQuery?

Yes, or really, the only thing this module does is call setString, setInteger, etc. for you. It's up to the application to then call executeQuery, addBatch, etc. for maximum flexibility. For instance in ADBC it's used with addBatch/executeBatch:

https://github.com/apache/arrow-adbc/blob/2485d7c3da217a7190f86128d769a7d0445755ab/java/driver/jdbc/src/main/java/org/apache/arrow/adbc/driver/jdbc/JdbcStatement.java#L160-L166

Cool! I believe this is a super useful feature. I'd like to review it in the following days.

Comment threaddocs/source/java/jdbc.rst Outdated
Currently, it is not possible to define a custom type conversion for a
supported or unsupported type.
Currently, it is not possible to override the type conversion for a
supported type, or define a new conversion for an unsupported type.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What I mean is that you can't implement a custom Consumer and have it be used, all you can do is change what type is assumed by the existing converters. But I'll clarify this

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks.

* \(1) Strings longer than Integer.MAX_VALUE bytes (the maximum length
of a Java ``byte[]``) will cause a runtime exception.
* \(2) If the timestamp has a timezone, the JDBC type defaults to
TIMESTAMP_WITH_TIMEZONE. If the timestamp has no timezone,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what would happen when a timezone is absent, the program would thrown an exception?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It'll just call setTimestamp(int, Timestamp) instead of setTimestamp(int, Timestamp, Calendar), I'll update the doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification.

*/
public abstract class BaseColumnBinder<V extends FieldVector> implements ColumnBinder {
protected V vector;
protected int jdbcType;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we declare the fields as final?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for catching that.

}

public BigIntBinder(BigIntVector vector, int jdbcType) {
super(vector, jdbcType);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is a type other than Types.BIGINT allowed here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, I wanted to allow things like binding an Int64 vector to an Int field, maybe that is too much flexibility though.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see. Thanks for the clarification.

return jdbcType == null ? new TinyIntBinder((TinyIntVector) vector) :
new TinyIntBinder((TinyIntVector) vector, jdbcType);
} else {
throw new UnsupportedOperationException(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we can extract this statement for all type widths, at the beginning of this method?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment on lines +39 to +40
JdbcParameterBinder(
final PreparedStatement statement,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this intentionally package private instead of private? Maybe add a comment on the relationship between the last two parameters?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, changed it to private, and added some docstrings + an explicit Preconditions check for the last two parameters.

@lidavidm

Copy link
Copy Markdown
MemberAuthor

I've let this sit for a while so having fixed an additional bug I found, I'll merge this now (though not for 9.0.0)

@lidavidmlidavidm changed the title ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters ARROW-17004: [Java] Add utility to bind Arrow data to JDBC parameters Jul 26, 2022
@lidavidm
lidavidm merged commit a5a2837 into apache:masterJul 26, 2022
@lidavidm
lidavidm deleted the arrow-17004 branch July 26, 2022 16:59
@github-actions

Copy link
Copy Markdown

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = bbf249e and contender = a5a2837. a5a2837 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Failed ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.36% ⬆️0.0%] ursa-thinkcentre-m75q
Buildkite builds:
[Failed] a5a28377 ec2-t3-xlarge-us-east-2
[Failed] a5a28377 test-mac-arm
[Finished] a5a28377 ursa-i9-9960x
[Finished] a5a28377 ursa-thinkcentre-m75q
[Failed] bbf249e0 ec2-t3-xlarge-us-east-2
[Failed] bbf249e0 test-mac-arm
[Finished] bbf249e0 ursa-i9-9960x
[Finished] bbf249e0 ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…apache#13589)
This extends the arrow-jdbc adapter to also allow taking Arrow data and using it to bind JDBC PreparedStatement parameters, allowing you to "round trip" data to a certain extent. This was factored out of arrow-adbc since it's not strictly tied to ADBC. Authored-by: David Li <li.davidm96@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@lidavidm@lwhite1@pitrou@liyafan82@ursabot@emkornfield