How to consume a single buffer & connection to array interchange #39

Description

@rgommers

For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
""" """ifcol.offset!=0:
raiseNotImplementedError("column.offset > 0 not handled yet")
ifcol.describe_nullnotin (0, 1):
raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
# Handle the dtype_dtype=col.dtypekind=_dtype[0]
bitwidth=_dtype[1]
if_dtype[0] notin (0, 1, 2, 20):
raiseRuntimeError("Not a boolean, integer or floating-point dtype")
_ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
_uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
_floats= {32: np.float32, 64: np.float64}
_np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
column_dtype=_np_dtypes[kind][bitwidth]
# No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
# NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
shape=(_buffer.bufsize// (bitwidth//8),))
returnx

From #38 (review) (@kkraus14 & @rgommers):

In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

__cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

From #38 (comment) (@jorisvandenbossche):

I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

Alternative/extension to the current design

We could change the plain memory description + __dlpack__ to:

  1. Implementations MUST support a memory description with ptr, bufsize, and device
  2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
  3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

(1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

What is missing for dealing with memory buffers

A summary of why this is hard is:

  1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
  2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

So what we are aiming for (ambitiously) is:

  • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
  • In pure Python

The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      How to consume a single buffer & connection to array interchange #39

      Description

      @rgommers

      For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

      For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

      The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

      defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
      """ """ifcol.offset!=0:
      raiseNotImplementedError("column.offset > 0 not handled yet")
      ifcol.describe_nullnotin (0, 1):
      raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
      # Handle the dtype_dtype=col.dtypekind=_dtype[0]
      bitwidth=_dtype[1]
      if_dtype[0] notin (0, 1, 2, 20):
      raiseRuntimeError("Not a boolean, integer or floating-point dtype")
      _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
      _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
      _floats= {32: np.float32, 64: np.float64}
      _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
      column_dtype=_np_dtypes[kind][bitwidth]
      # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
      ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
      data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
      # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
      shape=(_buffer.bufsize// (bitwidth//8),))
      returnx

      From #38 (review) (@kkraus14 & @rgommers):

      In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

      Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

      __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

      I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

      Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

      From #38 (comment) (@jorisvandenbossche):

      I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

      Alternative/extension to the current design

      We could change the plain memory description + __dlpack__ to:

      1. Implementations MUST support a memory description with ptr, bufsize, and device
      2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
      3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

      (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

      The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

      And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

      What is missing for dealing with memory buffers

      A summary of why this is hard is:

      1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
      2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

      So what we are aiming for (ambitiously) is:

      • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
      • In pure Python

      The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

      Metadata

      Metadata

      Assignees

      No one assigned

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          How to consume a single buffer & connection to array interchange #39

          Description

          @rgommers

          For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

          For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

          The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

          defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
          """ """ifcol.offset!=0:
          raiseNotImplementedError("column.offset > 0 not handled yet")
          ifcol.describe_nullnotin (0, 1):
          raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
          # Handle the dtype_dtype=col.dtypekind=_dtype[0]
          bitwidth=_dtype[1]
          if_dtype[0] notin (0, 1, 2, 20):
          raiseRuntimeError("Not a boolean, integer or floating-point dtype")
          _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
          _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
          _floats= {32: np.float32, 64: np.float64}
          _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
          column_dtype=_np_dtypes[kind][bitwidth]
          # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
          ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
          data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
          # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
          shape=(_buffer.bufsize// (bitwidth//8),))
          returnx

          From #38 (review) (@kkraus14 & @rgommers):

          In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

          Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

          __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

          I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

          Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

          From #38 (comment) (@jorisvandenbossche):

          I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

          Alternative/extension to the current design

          We could change the plain memory description + __dlpack__ to:

          1. Implementations MUST support a memory description with ptr, bufsize, and device
          2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
          3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

          (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

          The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

          And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

          What is missing for dealing with memory buffers

          A summary of why this is hard is:

          1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
          2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

          So what we are aiming for (ambitiously) is:

          • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
          • In pure Python

          The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

          Metadata

          Metadata

          Assignees

          No one assigned

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              How to consume a single buffer & connection to array interchange #39

              Description

              @rgommers

              For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

              For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

              The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

              defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
              """ """ifcol.offset!=0:
              raiseNotImplementedError("column.offset > 0 not handled yet")
              ifcol.describe_nullnotin (0, 1):
              raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
              # Handle the dtype_dtype=col.dtypekind=_dtype[0]
              bitwidth=_dtype[1]
              if_dtype[0] notin (0, 1, 2, 20):
              raiseRuntimeError("Not a boolean, integer or floating-point dtype")
              _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
              _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
              _floats= {32: np.float32, 64: np.float64}
              _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
              column_dtype=_np_dtypes[kind][bitwidth]
              # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
              ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
              data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
              # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
              shape=(_buffer.bufsize// (bitwidth//8),))
              returnx

              From #38 (review) (@kkraus14 & @rgommers):

              In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

              Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

              __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

              I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

              Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

              From #38 (comment) (@jorisvandenbossche):

              I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

              Alternative/extension to the current design

              We could change the plain memory description + __dlpack__ to:

              1. Implementations MUST support a memory description with ptr, bufsize, and device
              2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
              3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

              (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

              The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

              And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

              What is missing for dealing with memory buffers

              A summary of why this is hard is:

              1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
              2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

              So what we are aiming for (ambitiously) is:

              • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
              • In pure Python

              The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

              Metadata

              Metadata

              Assignees

              No one assigned

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  How to consume a single buffer & connection to array interchange #39

                  Description

                  @rgommers

                  For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

                  For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

                  The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

                  defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
                  """ """ifcol.offset!=0:
                  raiseNotImplementedError("column.offset > 0 not handled yet")
                  ifcol.describe_nullnotin (0, 1):
                  raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
                  # Handle the dtype_dtype=col.dtypekind=_dtype[0]
                  bitwidth=_dtype[1]
                  if_dtype[0] notin (0, 1, 2, 20):
                  raiseRuntimeError("Not a boolean, integer or floating-point dtype")
                  _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
                  _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
                  _floats= {32: np.float32, 64: np.float64}
                  _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
                  column_dtype=_np_dtypes[kind][bitwidth]
                  # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
                  ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
                  data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
                  # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
                  shape=(_buffer.bufsize// (bitwidth//8),))
                  returnx

                  From #38 (review) (@kkraus14 & @rgommers):

                  In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

                  Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

                  __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

                  I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

                  Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

                  From #38 (comment) (@jorisvandenbossche):

                  I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

                  Alternative/extension to the current design

                  We could change the plain memory description + __dlpack__ to:

                  1. Implementations MUST support a memory description with ptr, bufsize, and device
                  2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
                  3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

                  (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

                  The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

                  And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

                  What is missing for dealing with memory buffers

                  A summary of why this is hard is:

                  1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
                  2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

                  So what we are aiming for (ambitiously) is:

                  • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
                  • In pure Python

                  The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      How to consume a single buffer & connection to array interchange #39

                      Description

                      @rgommers

                      For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

                      For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

                      The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

                      defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
                      """ """ifcol.offset!=0:
                      raiseNotImplementedError("column.offset > 0 not handled yet")
                      ifcol.describe_nullnotin (0, 1):
                      raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
                      # Handle the dtype_dtype=col.dtypekind=_dtype[0]
                      bitwidth=_dtype[1]
                      if_dtype[0] notin (0, 1, 2, 20):
                      raiseRuntimeError("Not a boolean, integer or floating-point dtype")
                      _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
                      _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
                      _floats= {32: np.float32, 64: np.float64}
                      _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
                      column_dtype=_np_dtypes[kind][bitwidth]
                      # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
                      ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
                      data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
                      # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
                      shape=(_buffer.bufsize// (bitwidth//8),))
                      returnx

                      From #38 (review) (@kkraus14 & @rgommers):

                      In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

                      Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

                      __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

                      I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

                      Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

                      From #38 (comment) (@jorisvandenbossche):

                      I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

                      Alternative/extension to the current design

                      We could change the plain memory description + __dlpack__ to:

                      1. Implementations MUST support a memory description with ptr, bufsize, and device
                      2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
                      3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

                      (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

                      The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

                      And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

                      What is missing for dealing with memory buffers

                      A summary of why this is hard is:

                      1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
                      2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

                      So what we are aiming for (ambitiously) is:

                      • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
                      • In pure Python

                      The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          How to consume a single buffer & connection to array interchange #39

                          Description

                          @rgommers

                          For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

                          For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

                          The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

                          defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
                          """ """ifcol.offset!=0:
                          raiseNotImplementedError("column.offset > 0 not handled yet")
                          ifcol.describe_nullnotin (0, 1):
                          raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
                          # Handle the dtype_dtype=col.dtypekind=_dtype[0]
                          bitwidth=_dtype[1]
                          if_dtype[0] notin (0, 1, 2, 20):
                          raiseRuntimeError("Not a boolean, integer or floating-point dtype")
                          _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
                          _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
                          _floats= {32: np.float32, 64: np.float64}
                          _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
                          column_dtype=_np_dtypes[kind][bitwidth]
                          # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
                          ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
                          data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
                          # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
                          shape=(_buffer.bufsize// (bitwidth//8),))
                          returnx

                          From #38 (review) (@kkraus14 & @rgommers):

                          In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

                          Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

                          __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

                          I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

                          Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

                          From #38 (comment) (@jorisvandenbossche):

                          I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

                          Alternative/extension to the current design

                          We could change the plain memory description + __dlpack__ to:

                          1. Implementations MUST support a memory description with ptr, bufsize, and device
                          2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
                          3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

                          (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

                          The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

                          And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

                          What is missing for dealing with memory buffers

                          A summary of why this is hard is:

                          1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
                          2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

                          So what we are aiming for (ambitiously) is:

                          • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
                          • In pure Python

                          The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              How to consume a single buffer & connection to array interchange #39

                              Description

                              @rgommers

                              For dataframe interchange, the smallest building block is a "buffer" (see gh-35, gh-38) - a block of memory. Interpreting that is nontrivial, especially if the goal is to build an interchange protocol in Python. That's why DLPack, buffer protocol, __array_interface__, __cuda_array_interface__, __array__ and __arrow_array__ all exist, and are still complicated.

                              For what a buffer is, currently it's only a data pointer (ptr) and a size (bufsize) which together describe a contiguous block of memory, plus a device attribute (__dlpack_device__) and optionally DLPack support (__dlpack__). One open question is:

                              The other, larger question is how to make buffers nice to deal with for implementers of the protocol. The current Pandas prototype shows the issue:

                              defconvert_column_to_ndarray(col : ColumnObject) ->np.ndarray:
                              """ """ifcol.offset!=0:
                              raiseNotImplementedError("column.offset > 0 not handled yet")
                              ifcol.describe_nullnotin (0, 1):
                              raiseNotImplementedError("Null values represented as masks or ""sentinel values not handled yet")
                              # Handle the dtype_dtype=col.dtypekind=_dtype[0]
                              bitwidth=_dtype[1]
                              if_dtype[0] notin (0, 1, 2, 20):
                              raiseRuntimeError("Not a boolean, integer or floating-point dtype")
                              _ints= {8: np.int8, 16: np.int16, 32: np.int32, 64: np.int64}
                              _uints= {8: np.uint8, 16: np.uint16, 32: np.uint32, 64: np.uint64}
                              _floats= {32: np.float32, 64: np.float64}
                              _np_dtypes= {0: _ints, 1: _uints, 2: _floats, 20: {8: bool}}
                              column_dtype=_np_dtypes[kind][bitwidth]
                              # No DLPack yet, so need to construct a new ndarray from the data pointer# and size in the buffer plus the dtype on the column_buffer=col.get_data_buffer()
                              ctypes_type=np.ctypeslib.as_ctypes_type(column_dtype)
                              data_pointer=ctypes.cast(_buffer.ptr, ctypes.POINTER(ctypes_type))
                              # NOTE: `x` does not own its memory, so the caller of this function must# either make a copy or hold on to a reference of the column or# buffer! (not done yet, this is pretty awful ...)x=np.ctypeslib.as_array(data_pointer,
                              shape=(_buffer.bufsize// (bitwidth//8),))
                              returnx

                              From #38 (review) (@kkraus14 & @rgommers):

                              In __cuda_array_interface__ we've generally stated that holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well.

                              Yes that works and I've thought about it. The trouble is where to hold the reference. You really need one reference per buffer, not just store a reference to the whole exchange dataframe object (buffers can end up elsewhere outside the new pandas dataframe here). And given that a buffer just has a raw pointer plus a size, there's nothing to hold on to. I don't think there's a sane pure Python solution.

                              __cuda_array_interface__ is directly attached to the object you need to hold on to, which is not the case for this Buffer.

                              I'd argue this is a place where we should really align with the array interchange protocol though as the same problem is being solved there.

                              Yep, for numerical data types the solution can simply be: hurry up with implementing __dlpack__, and the problem goes away. The dtypes that DLPack does not support are more of an issue.

                              From #38 (comment) (@jorisvandenbossche):

                              I personally think it would be useful to keep those existing interface methods (or array, or arrow_array). For people that are using those interface, that will be easier to interface with the interchange protocol than manually converting the buffers.

                              Alternative/extension to the current design

                              We could change the plain memory description + __dlpack__ to:

                              1. Implementations MUST support a memory description with ptr, bufsize, and device
                              2. Implementations MAY support buffers in their native format (e.g. add a native enum attribute, and if both producer and consumer happen to use that native format, they can call the corresponding protocol - __arrow_array__ or __array__)
                              3. Implementations MAY support any exchange protocol (DLPack, __cuda_array_interface__, buffer protocol, __array_interface__).

                              (1) is required for any implementation to be able to talk to any other implementation, but also the most clunky to support because it needs to solve the "who owns this memory and how do you prevent it from being freed" all over again. What is needed there is

                              The advantage of (2) and (3) are that they have the most hairy issue already solved, and will likely be faster.

                              And the MUST/MAY should address @kkraus14's concern that people will just standardize on the lowest common denominator (numpy).

                              What is missing for dealing with memory buffers

                              A summary of why this is hard is:

                              1. Underlying implementations are not compatible. E.g., NumPy doesn't support variable length strings or bit masks, Arrow does not support strided arrays or byte masks.
                              2. DLPack is the only protocol with device support, but it does not support all dtypes that are needed.

                              So what we are aiming for (ambitiously) is:

                              • Something flexible enough to be a superset of NumPy and Arrow, with full device support.
                              • In pure Python

                              The "holding a reference to the producing object must guarantee the lifetime of the memory and that has worked relatively well" seems necessary for supporting the raw memory description. This probably means that (a) the Buffer object should include the right Python object to keep a reference to (for Pandas that would typically be a 1-D numpy array), and (b) there must be some machinery to keep this reference alive (TBD what that looks like, likely not pure Python) in the implementation.

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions