Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Read and bind a byte string - #38

Merged
tamnd merged 1 commit into
mainfrom
bytes
Aug 24, 2026
Merged

Read and bind a byte string#38
tamnd merged 1 commit into
mainfrom
bytes

Conversation

@tamnd

Copy link
Copy Markdown
Owner

Closes#37.

The issue said the conformance reader was the only thing missing, and that was true of the reader and not of the client. The pin was at 6ee7d80, which is before the engine had a Value::Bytes at all, so there was nothing for to_py to return and nothing for a parameter to become. The pin moves to 230581c and the three places that touch a value gain the arm they were missing.

Reading gives bytes and binding takes bytes. Not bytearray and not memoryview: a value that came out of a result is a reading of what the file holds and nothing in Python should be able to write through it, and a parameter is read after the call that takes it returns, so a buffer the caller can still write through is a promise this client would be taking on trust. One type in all three directions, which is the one the loader already names for a byte string column.

numpy gets an object array. A byte string column keeps the two buffers a string column keeps, so the walk is the string walk without the UTF-8 check, but the array cannot be an S one: S pads every cell to the longest and drops trailing nulls, which is a different value from the one stored.

The reader turns hexits into octets itself rather than calling bytes.fromhex. The two agree on everything except a vertical tab, which fromhex drops and Rust's is_ascii_whitespace does not, and which whitespace fromhex drops has changed across the Python versions this client supports. A reader of the shared corpus that accepts a shade more than the reference one is a reader that lets a malformed case through on one client and not on another, which is the failure the encoding exists to make impossible.

The issue asked whether the comparison path keeps a byte string apart from the string that spells the same octets, since the case at string.yaml line 995 asserts that X'0041' = 'A' is not true and a reader that decoded octets into a str somewhere would pass it for the wrong reason. It does, and there is now a test saying so on both sides of the comparison.

The pin bump also made ten words reserved that were not before: on, at, number, nothing, record, small, count, big, day and exact. Four test files used four of them as an alias or a column name and are renamed. Nothing in the client changed for that, only what the tests are allowed to call things. Every other client will hit the same list on its next pin bump.

A byte string can be stored and cannot be read back yet, so the loader's refusal at src/load.rs:148 stays where it is. ColType has no byte string in it and the row walk has no arm that produces one, which is tamnd/zu#728.

What was run

Built and tested on a Linux box against CPython 3.12.3.

The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all six are a time written to the nanosecond, which is a digit finer than a Python datetime holds. Before this the six BYTES cases in string.yaml were refused at load rather than run.

The six test files this touches are 360 passed and 2 skipped, the two skips being the corpus tests that want ZU_CASES. ruff check, ruff format and cargo fmt are clean.

The issue said the conformance reader was the only thing missing, and
that was true of the reader and not of the client. The pin was at
6ee7d80, which is before the engine had a Value::Bytes at all, so there
was nothing for to_py to return and nothing for a parameter to become.
The pin moves to 230581c and the three places that touch a value gain
the arm they were missing.
Reading gives bytes and binding takes bytes. Not bytearray and not
memoryview: a value that came out of a result is a reading of what the
file holds and nothing in Python should be able to write through it,
and a parameter is read after the call that takes it returns, so a
buffer the caller can still write through is a promise this client
would be taking on trust. One type in all three directions, which is
the one the loader already names for a byte string column.
numpy gets an object array. A byte string column keeps the two buffers
a string column keeps, so the walk is the string walk without the UTF-8
check, but the array cannot be an S one: S pads every cell to the
longest and drops trailing nulls, which is a different value from the
one stored.
The reader turns hexits into octets itself rather than calling
bytes.fromhex. The two agree on everything except a vertical tab, which
fromhex drops and Rust's is_ascii_whitespace does not, and which
whitespace fromhex drops has changed across the Python versions this
client supports. A reader of the shared corpus that accepts a shade
more than the reference one is a reader that lets a malformed case
through on one client and not on another, which is the failure the
encoding exists to make impossible.
The pin bump also made ten words reserved that were not before: on, at,
number, nothing, record, small, count, big, day and exact. Four test
files used four of them as an alias or a column name and are renamed.
Nothing in the client changed for that, only what the tests are allowed
to call things.
A byte string can be stored and cannot be read back yet, so the
loader's refusal at src/load.rs:148 stays where it is. ColType has no
byte string in it and the row walk has no arm that produces one, which
is tamnd/zu#728.
The corpus is 1399 cases, 1393 passed, 0 failed, 6 unsupported, and all
six are a time written to the nanosecond, which is a digit finer than a
Python datetime holds. The six test files this touches are 360 passed
and 2 skipped, the two skips being the corpus tests that want ZU_CASES.
@tamnd
tamnd merged commit a8a8330 into mainAug 24, 2026
29 of 33 checks passed
@tamnd
tamnd deleted the bytes branch August 24, 2026 21:51
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The runner calls BYTES reserved, so six corpus cases refuse to load

1 participant

@tamnd