ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760) - #13811

Merged
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master
Aug 8, 2022
Merged

ARROW-17303: [Java][Dataset] Read Arrow IPC files by NativeDatasetFactory (#13760)#13811
lidavidm merged 5 commits into
apache:masterfrom
igor-suhorukov:master

Conversation

@igor-suhorukov

@igor-suhorukovigor-suhorukov commented Aug 7, 2022

Copy link
Copy Markdown
Contributor

This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(), FileFormat.ARROW_IPC, arrowDatasetURL);

It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion

@github-actions

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown

⚠️ Ticket has not been started in JIRA, please click 'Start Progress'.

@lidavidm

Copy link
Copy Markdown
Member

Thanks for the PR!

@davisusanibar@lwhite1 would one of you mind taking a look?

Is "osm_nodes.arrow" from OpenStreetMap? Are there licensing concerns around the data? Arrow already has test data files for use and/or files can be generated in-process.

@igor-suhorukov

igor-suhorukov commented Aug 8, 2022

Copy link
Copy Markdown
ContributorAuthor

@lidavidm yes, it is 10 records from Openstreetmap planet dump. Could you please provide more information how to generate test data in ARROW file format to test dataset API or where existing test data located?

@lidavidm

Copy link
Copy Markdown
Member

It'd be something like

Fileout = TMP.newFile();
Schemaschema = newSchema(Collections.singletonList(Field.nullable("ints", newArrowType.Int(32, true))));
try (VectorSchemaRootroot = VectorSchemaRoot.create(schema, allocator);
FileOutputStreamfileOutputStream = newFileOutputStream(file);
ArrowFileWriterwriter = newArrowFileWriter(root, /*dictionaryProvider=*/null, sink)) {
// Fill root with dataIntVectorints = (IntVector) root.getVector(0);
ints.setSafe(0, 0);
root.setRowCount(1);
// ...writer.start();
writer.writeBatch();
writer.end();
}
// Use out.getPath()...

@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

@lidavidm thank you for advise. OSM data was deleted from PR. Please check updated test TestFileSystemDataset#testBaseArrowIpcRead
Is it fit project test approach?

@lwhite1

Copy link
Copy Markdown
Contributor

Hi @igor-suhorukov This looks good to me except I wish the tests were more robust. (The same is true for the Parquet test that you're emulating, but I guess that's out of scope here.)

This kind of test - relying on checking sizes and names - doesn't provide much assurance that we won't see bug reports when people import complex data types or otherwise tap into some of the more advanced functionality.

@lidavidm

Copy link
Copy Markdown
Member

@lwhite1 we could file another JIRA for that?

@lidavidm

Copy link
Copy Markdown
Member

Also a general note re: Larry's comment: we currently have a mix of JUnit 4/5, ad-hoc test helpers like the one here, and a mix of assertion libraries; it might be good to start incrementally cleaning that up (e.g. it would be much easier to test complex types if there were an easy setup to parameterize a test and have the data generated for you).

ARROW-6931 is sort of related, and ARROW-4740 (we added JUnit5 but didn't port the existing tests)

@lwhite1

lwhite1 commented Aug 8, 2022 via email

Copy link
Copy Markdown
Contributor

@lidavidm

Copy link
Copy Markdown
Member

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

@lwhite1

Copy link
Copy Markdown
Contributor

I suggest a separate ticket because 1) generating test data is very unergonomic (as seen here) and could use some thought across different areas of the codebase and 2) I'd rather push down the testing to the appropriate levels (IPC, Parquet, and eventually CSV should share most of their testing code, the same way the C++ library is organized; and most of the type-specific tests should be done for the C Data Interface)

Ok. Works for me.

@lwhite1lwhite1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@lidavidm

Copy link
Copy Markdown
Member

I filed https://issues.apache.org/jira/browse/ARROW-17342

@lidavidm

Copy link
Copy Markdown
Member

FWIW, looking at the JIRA/GH issue, this will only handle "IPC" files, not Arrow stream files - there's work needed on the C++ side if that is something we want to cover

@lidavidm
lidavidm merged commit 78351ce into apache:masterAug 8, 2022
@igor-suhorukov

Copy link
Copy Markdown
ContributorAuthor

Thanks a lot for clarification @lidavidm@lwhite1 and for your time. Don't worries about refactoring. I have such experience with Spring/ElasticSearch projects refactoring, fix tech debt and cleanup - it can be contribution of crowd when Arrow project will be more mature - separate activities for new joiners. Good start for someone

@ursabot

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = a2f3666 and contender = 78351ce. 78351ce is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Finished ⬇️0.34% ⬆️0.0%] test-mac-arm
[Finished ⬇️0.0% ⬆️0.0%] ursa-i9-9960x
[Finished ⬇️0.14% ⬆️0.04%] ursa-thinkcentre-m75q
Buildkite builds:
[Finished] 78351cec ec2-t3-xlarge-us-east-2
[Finished] 78351cec test-mac-arm
[Finished] 78351cec ursa-i9-9960x
[Finished] 78351cec ursa-thinkcentre-m75q
[Finished] a2f3666d ec2-t3-xlarge-us-east-2
[Finished] a2f3666d test-mac-arm
[Finished] a2f3666d ursa-i9-9960x
[Finished] a2f3666d ursa-thinkcentre-m75q
Supported benchmarks:
ec2-t3-xlarge-us-east-2: Supported benchmark langs: Python, R. Runs only benchmarks with cloud = True
test-mac-arm: Supported benchmark langs: C++, Python, R
ursa-i9-9960x: Supported benchmark langs: Python, R, JavaScript
ursa-thinkcentre-m75q: Supported benchmark langs: C++, Java

Yicong-Huang added a commit to apache/texera that referenced this pull request Dec 13, 2022
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
pribor pushed a commit to GlobalWebIndex/arrow that referenced this pull request Oct 24, 2025
…tory (apache#13760) (apache#13811)
This PR allow developers to create Dataset from ARROW IPC files in JVM code like:
`FileSystemDatasetFactory factory = new FileSystemDatasetFactory(rootAllocator(), NativeMemoryPool.getDefault(),
FileFormat.ARROW_IPC, arrowDatasetURL);`
It is foundation for Apache Spark arrow data source to process huge existing partitioned datasets in ARROW file format without additional data format conversion
Lead-authored-by: Igor Suhorukov <igor.suhorukov@gmail.com>
Co-authored-by: igor.suhorukov <igor.suhorukov@gmail.com>
Signed-off-by: David Li <li.davidm96@gmail.com>
yangzhang75 pushed a commit to yangzhang75/texera that referenced this pull request Jun 22, 2026
This PR bumps Apache Arrow version from 9.0.0 to 10.0.0.
Main changes related to PyAmber:
## Java/Scala side:
- JDBC Driver for Arrow Flight SQL
([13800](apache/arrow#13800))
- Initial implementation of immutable Table API
([14316](apache/arrow#14316))
- Substrait, transaction, cancellation for Flight SQL
([13492](apache/arrow#13492))
- Read Arrow IPC, CSV, and ORC files by NativeDatasetFactory
([13811](apache/arrow#13811),
[13973](apache/arrow#13973),
[14182](apache/arrow#14182))
- Add utility to bind Arrow data to JDBC parameters
([13589](apache/arrow#13589))
## Python side:
- The batch_readahead and fragment_readahead arguments for scanning
Datasets are exposed in Python
([ARROW-17299](https://issues.apache.org/jira/browse/ARROW-17299)).
- ExtensionArrays can now be created from a storage array through the
pa.array(..) constructor
([ARROW-17834](https://issues.apache.org/jira/browse/ARROW-17834)).
- Converting ListArrays containing ExtensionArray values to numpy or
pandas works by falling back to the storage array
([ARROW-17813](https://issues.apache.org/jira/browse/ARROW-17813)).
- Casting Tables to a new schema now honors the nullability flag in the
target schema
([ARROW-16651](https://issues.apache.org/jira/browse/ARROW-16651)).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read "arrow" (IPC and streaming) files usning org.apache.arrow.dataset.jni.NativeDatasetFactory in Java API

4 participants

@igor-suhorukov@lidavidm@lwhite1@ursabot