A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

A result as numpy arrays - #30

Merged
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy
Aug 19, 2026
Merged

A result as numpy arrays#30
tamnd merged 1 commit into
mainfrom
dx-fetchnumpy

Conversation

@tamnd

Copy link
Copy Markdown
Owner

result.fetchnumpy(), which is a dict of numpy arrays keyed by column name, and it needs numpy and nothing else.

It exists because numpy is what a lot of code already holds: a model that takes arrays, a plotting call, a loop somebody wrote before DataFrames. Handing that code an Arrow table means it has to convert one, and converting one means pyarrow has to be installed to do work numpy could have done with the same bytes.

The same bytes is what this is. zudb::query::column already hands back one owned buffer per column, values end to end in the layout every columnar format uses, and a numpy array is a pointer, a length and a dtype over exactly that. So an integer, float, datetime or duration column becomes an array by moving the Vec into numpy and naming the type: no pass over the values, no allocation, no second copy of the result. There is a test for the claim, since it is the only one worth making: owndata is false and there is a base object, which is numpy's own way of saying the memory came from somewhere else.

Two columns cost a pass because the layouts differ rather than the names. A boolean is a bit per row here and a byte per row there. A date is 32 bits here and 64 in a datetime64. Both are one widening pass and neither has another way.

The mapping, and each line of it has a test:

columnnumpy
INT, FLOAT, BOOLint64, float64, bool
DATEdatetime64[D]
DATETIME, ZONED DATETIMEdatetime64[ns], the instant in UTC
TIME, day-time DURATIONtimedelta64[ns]
year-month DURATIONtimedelta64[M]
STRING, node, rel, path, list, recordobject array

A time of day is nanoseconds since midnight, which is what a clock reading is on a number line and the closest thing numpy has to one. A datetime with an offset is the instant, because numpy has no zone to keep the offset in. A time with an offset is refused, which is the same refusal the Arrow path makes and for the same reason: dropping the offset would move the value.

Nulls are the other half of the shape. numpy has no missing integer, so a column with a null in it comes back as a numpy.ma.masked_array: the data array underneath is still the engine's buffer and the mask is built from the validity bitmap the engine already filled, so masking is a wrapper and not a copy. A column with nothing missing is a plain array, because a mask nothing is masked by is a second buffer nobody asked for. Object columns carry None in the cell instead, since they have somewhere to put one.

Two columns of the same name are refused rather than silently collapsed into one dict entry, and the message says to name them apart.

Over a million rows in two columns on this machine: fetchnumpy() 26 ms, against 47 ms for building the Arrow table and calling to_numpy on each of its columns, and 130 ms for the same rows as tuples.

The extension links numpy's C API through rust-numpy, which loads it at the first array rather than at module init, so import zudb still imports numpy no more than it imports pandas. There is a test that says so and a line in the clean-machine smoke test that checks the refusal reads pip install 'zudb[numpy]'.

27 tests in tests/test_numpy.py, the stub, the numpy extra in pyproject, and a paragraph in the columns section of the README. Green locally: ruff, ruff format, clippy, cargo fmt, and the suite.

@tamnd
tamnd merged commit a022c14 into mainAug 19, 2026
29 of 35 checks passed
@tamnd
tamnd deleted the dx-fetchnumpy branch August 19, 2026 12:48
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@tamnd