ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory - #14182

Merged
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786
Oct 11, 2022
Merged

ARROW-17786: [Java] Read CSV files using org.apache.arrow.dataset.jni.NativeDatasetFactory#14182
lidavidm merged 7 commits into
apache:masterfrom
davisusanibar:ARROW-17786

Conversation

@davisusanibar

Copy link
Copy Markdown
Contributor

Support CSV file format in java Dataset API

@github-actions

Copy link
Copy Markdown


public CsvWriteSupport(File outputFolder) {
path = outputFolder.getPath() + File.separator + "generated-" + random.nextLong() + ".csv";
uri = "file://" + path;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, changed

@pitrou

Copy link
Copy Markdown
Member

@lwhite1 Would you like to review this?

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

@pitrou

Copy link
Copy Markdown
Member

@davisusanibar Why is only write support tested? Shouldn't you test support for dataset reads as well?

Hi, write support is doing by Java native libraries, then Dataset module is reading that CSV file with (FileFormat.CSV): new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.CSV, writeSupport.getOutputURI());

Currently we are only offering Read support for Parquet. ORC and now CSV, ... write support it not implemented

Oops, sorry, I had misread the tests.

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Related to update Java Dataset module documentation, this Jira will cover that: https://issues.apache.org/jira/browse/ARROW-17789

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {
was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

This line

if (ARROW_PREDICT_TRUE(SetFieldBuilder(string_view(key, len), &duplicate_keys))) {

was not changed long time ago, I don't know what is the error:

[1/1] cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
FAILED: CMakeFiles/lint cd /tmp/arrow-lint-5g90mrbe/cpp-build && /usr/local/bin/python /arrow/cpp/build-support/run_cpplint.py --cpplint_binary /arrow/cpp/build-support/cpplint.py --exclude_globs /arrow/cpp/build-support/lint_exclusions.txt --source_dir /arrow/cpp/src --source_dir /arrow/cpp/examples --source_dir /arrow/cpp/tools --quiet
/arrow/cpp/src/arrow/json/parser.cc:1030: Lines should be <= 90 characters long [whitespace/line_length] [2]
Done processing /arrow/cpp/src/arrow/json/parser.cc
Total errors found: 1

Just rebase and working fine now

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please, when you have some time if you could review this, thank you

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Please if you could help me with a review @lwhite1@lidavidm

Comment threadjava/dataset/src/main/cpp/jni_wrapper.cc

checkParquetReadResult(schema, expectedJsonUnordered, datum);

AutoCloseables.close(datum);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use try-with-resources?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added try-with for FileSystemDatasetFactory but don't know how to add that forList<ArrowRecordBatch> that is an input for another function.

@lwhite1

lwhite1 commented Oct 6, 2022

Copy link
Copy Markdown
Contributor

Sorry for not getting to this sooner. Is this for release 10?

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

Sorry for not getting to this sooner. Is this for release 10?

Not problem for that. Yes, trying to merge for 10 to then could create cookbooks for this CSV Reader case

@lidavidm

Copy link
Copy Markdown
Member

@github-actions crossbow submit java

@github-actions

Copy link
Copy Markdown

Revision: 98c6d1d

Submitted crossbow builds: ursacomputing/crossbow @ actions-e54f92cb9c

TaskStatus
java-jarsGithub Actions
verify-rc-source-java-linux-almalinux-8-amd64Github Actions
verify-rc-source-java-linux-conda-latest-amd64Github Actions
verify-rc-source-java-linux-ubuntu-18.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-20.04-amd64Github Actions
verify-rc-source-java-linux-ubuntu-22.04-amd64Github Actions
verify-rc-source-java-macos-amd64Github Actions

@lidavidm

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lidavidm

lidavidm commented Oct 7, 2022

Copy link
Copy Markdown
Member

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

log link: https://github.com/apache/arrow/actions/runs/3200943136/jobs/5229379128 (I'm going to restart the job now)

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

@davisusanibar

Copy link
Copy Markdown
ContributorAuthor

@davisusanibar it may be worth investigating if we can not build/run tests in java-jars? It seems rather redundant (or maybe that's on purpose)

Yes, sound redundant at first review, but, let me check what part could be improved

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

@lwhite1

Copy link
Copy Markdown
Contributor

CI shows a flake in testTable:

Error: org.apache.arrow.c.RoundtripTest.testTable Time elapsed: 0.031 s <<< ERROR!
java.lang.IllegalStateException: Cannot import released ArrowSchema
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:458)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:63)
at org.apache.arrow.c.SchemaImporter.importField(SchemaImporter.java:56)
at org.apache.arrow.c.Data.importField(Data.java:246)
at org.apache.arrow.c.Data.importSchema(Data.java:266)
at org.apache.arrow.c.Data.importVectorSchemaRoot(Data.java:377)
at org.apache.arrow.c.RoundtripTest.testTable(RoundtripTest.java:683)

@lwhite1 have you seen this? Might be worth some investigation

I haven't noticed it before but there are so many failures with any given commit I probably don't pay as much attention as I should to something that doesn't look related to the changes. I will take a look now.

It looks like a bug in the test. I will push a fix. IDK why it passes locally.

@lidavidm

Copy link
Copy Markdown
Member

It looks like the test crashes on another attempt.

@lidavidm

Copy link
Copy Markdown
Member

@davisusanibar can you rebase here to see if Larry's fix makes CI pass?

@lidavidm

Copy link
Copy Markdown
Member

Ok, looks all good. Thanks!

@lidavidm
lidavidm merged commit a39f219 into apache:masterOct 11, 2022
@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = fa3cf78 and contender = a39f219. a39f219 is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed ⬇️1.67% ⬆️0.0%] test-mac-arm
[Failed ⬇️0.27% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] a39f2197 ec2-t3-xlarge-us-east-2
[Failed] a39f2197 test-mac-arm
[Failed] a39f2197 ursa-i9-9960x
[Finished] a39f2197 ursa-thinkcentre-m75q
[Finished] fa3cf78e ec2-t3-xlarge-us-east-2
[Failed] fa3cf78e test-mac-arm
[Failed] fa3cf78e ursa-i9-9960x
[Finished] fa3cf78e ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

@ursabot

Copy link
Copy Markdown

['Python', 'R'] benchmarks have high level of regressions.
ursa-i9-9960x

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
lidavidm pushed a commit that referenced this pull request Mar 2, 2023
…upport (as per ARROW-17786) (#34390)
Update status documentation: CSV read in Java from #14182
Authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
….NativeDatasetFactory (apache#14182)
Support CSV file format in java Dataset API
Authored-by: david dali susanibar arce <davi.sarces@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@davisusanibar@pitrou@lwhite1@lidavidm@ursabot