Uh oh!
There was an error while loading. Please reload this page.
Add variable-length string support - #45
Conversation
* Add a summary document for the dataframe interchange protocol Summarizes the various discussions about and goals/non-goals and requirements for the `__dataframe__` data interchange protocol. The intended audience for this document is Consortium members and dataframe library maintainers who may want to support this protocol. The aim is to keep updating this till we have captured all the requirements and answered all the FAQs, so we can actually design the protocol after and verify it meets all our requirements. Closesgh-29 * Process some review comments * Process a few more review comments. * Link to Release callback semantics in Arrow C Data Interface docs * Add design requirements for column selection and df metadata * Edit the nested/heterogeneous dtypes non-requirement * Add requirements for chunking and memory layout description Also address some smaller review comments. * Add TBD notes on dataframe-array connection and from_dataframe Also add more details on the Arrow C Data Interface. * Address review comments * Add details on implementation options * Add details about the C implementation * Add an image of the dataframe model and its memory layout. * Add link to discussion on array-dataframe connection * Some more updates for review comments * Update table to indicate Arrow does support categoricals. * Add section on dtype format strings * Reflow some lines * Add a requirement on semantic meaning of NaN/NaT, and timezone detail * Textual tweak: say columns in a data frame are ordered * Update requirements document for recent decisions/insights
Add a prototype of the dataframe interchange protocol
rgommers
commented
Jun 28, 2021
Thanks @kgryte! It may be useful to close this PR and resend it against |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
…o variable-length-string-support
kgryte
commented
Jun 28, 2021
@rgommers Will close this and submit against |
jorisvandenbossche
commented
Jul 8, 2021
You can actually change the target branch by clicking the "Edit" button next to title, so then you don't need to close / open a new PR |
Uh oh!
There was an error while loading. Please reload this page.
jorisvandenbossche
left a comment
There was a problem hiding this comment.
I think the separate get_data_buffer and get_offsets methods might be a bit problematic. Typically you only want to do one pass over the data and create all buffers at once.
I know this is only a dummy implementation, and presumably the created buffers could be cached on the object so the separate methods don't calculate it twice. But might be useful to think about separate methods vs a single get_buffers() (with a specified order)
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
kgryte
commented
Jul 8, 2021
I initially tried doing that, but doing so added many unrelated changes and muddied this PR. :( |
jorisvandenbossche
commented
Jul 8, 2021
I think that either merging latest master in this branch or rebasing this branch on top of master should solve that |
This is a fresh port of changes made in order to support variable length strings in order to provide a cleaner merge.
kgryte
commented
Jul 19, 2021
Closing this PR out in favor of gh-47. |
This PR
offsetsandmaskbuffers.objectdtype. The implementation will need to be updated to accommodate pandas' string extension dtype which is based on arrow. Currently, the string extension dtype is considered experimental and subject to change. The use ofobjectdtype is still used as the default string dtype for backward compatiblity.__dataframe__uses a bit array to indicate missing values.