Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Ensure DictStore contains only bytes - #350

Merged
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes
Dec 7, 2018
Merged

Ensure DictStore contains only bytes#350
alimanfoo merged 8 commits into
zarr-developers:masterfrom
jakirkham:ensure_DictStore_contains_only_bytes

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 3, 2018

Copy link
Copy Markdown
Member

Partially addresses issue ( #348 ) and issue ( #349 ).

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint. Of course nesting of DictStores are still allowed per usual. However these really just map to a variety of keys, which is fine.

Since the DictStore's values are just bytes, there shouldn't be any cases where the size of these values cannot be determined. So drop handling for unknown sizes in buffer_size. Also drop the associated test for DictStore as this cannot occur.

Add a test case for getsize with a non-conforming dict-based store where sizes are unknown to make sure that case is tested and handled appropriately.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Just for clarity, would be perfectly happy with any option that enforces that bytes-like data is in DictStore. Simply view this as a step in that direction. We can of course change to using memoryviews or anything else internally that makes sense and we are comfortable with.

@alimanfoo

Copy link
Copy Markdown
Member

As the spec notes stores values must be an "arbitrary sequence of bytes", this change ensures that values in DictStore meet that constraint.

Just to note, the spec really talks at an abstract level, so "arbitrary sequence of bytes" should be interpreted as an abstract concept. In a given programming language there may be many different types of object that encapsulate an arbitrary sequence of bytes, and the spec does not constrain which should be used. E.g., Python has the buffer protocol, which is a standard way of objects declaring that they encapsulate a sequence of bytes stored in memory. So in a Python implementation, it would be appropriate to require that stores can accept any value that exports the buffer protocol, but that does not have to be a bytes object.

That said, there are separate reasons for thinking that the DictStore should internally convert values to bytes objects, e.g., to ensure they are immutable. Sorry for being pedantic, but just wanted to clarify it is not the spec that requires that.

@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham. FWIW I'm happy to go ahead with this, but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

No worries. Completely agree.

By "meet that constraint" wasn't trying to suggest that the spec meant bytes specifically. Was more meaning that by using ensure_bytes we would reject anything that was not spec conforming from being stored. Tried to clarify this with the follow-up comment above. Sorry if that was still unclear.

Hope this makes more sense. Please let me know if I'm still missing something.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...but if we do I think we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes. Otherwise there will be a performance hit for in-memory zarr usage due to the extra memory copy.

Completely agree for the same reasons. Am working on that currently.

@alimanfoo

alimanfoo commented Dec 3, 2018 via email

Copy link
Copy Markdown
Member

@jakirkham

Copy link
Copy Markdown
MemberAuthor

...we should back out the changes in numcodecs that modified encode() methods to return ndarray when previously they returned bytes.

Am working on that currently.

PR ( zarr-developers/numcodecs#155 ) should handle this.

@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As the spec requires that the data in a store be a sequence of `bytes`,
make sure that non-`DictStore` input meets this requirement when setting
values. This effectively ensures that other `DictStore` meet this
requirement as well. So we don't need to go through and check their
values too.
As everything in `DictStore` must either be another `DictStore` or
`bytes`, there shouldn't be any cases where the size is undefined nor
cases that this exception should need handling. Given this go ahead and
drop the special casing for unknown sizes in `DictStore`.
While this test case does test a useful subset of the `getsize` API, the
contents being added to the store here are non-conforming to our
expectations of store contents. Namely the store should only contain
values that are an "arbitrary sequence of bytes", which this test case
is not.
This creates a non-conforming store to make sure that `getsize` handles
its contents in the expected way. Namely that it returns `-1`.
@jakirkham

jakirkham commented Dec 6, 2018

Copy link
Copy Markdown
MemberAuthor

Please let me know if there is anything else needed for this.

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, I'm happy for this to go in as-is. Couple of thoughts but up to your discretion whether to address in this PR:

  • Do you want to handle the renaming of DictStore to MemoryStore here? Or better in a separate PR?
  • Maybe worth a test to verify that trying to set a non-buffer-like value in a DictStore raises a TypeError?

@jakirkhamjakirkham mentioned this pull request Dec 6, 2018
Add a test to ensure that a non-buffer supporting object when stored
into a valid store, will raise a `TypeError` instead of storing it.
Disable this checking for generic `MappingStore`s (e.g. `dict`) as they
do not perform this sort of checking on the data they accept as values.
Provide a simple test for `DictStore` to ensure that non-`bytes` is
coerced to `bytes` before storing it and is retrieved as `bytes`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it. Though I agree it is a good idea.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Comment threadzarr/storage.py Outdated
@alimanfoo

Copy link
Copy Markdown
Member

IMHO opinion renaming the store is a separate topic. Have raised issue ( #356 ) to track it.

Great, thanks.

Adding some tests seems reasonable for this PR. Have included handling of the non-buffer case (disabled for dict). Also have added a test to ensure that DictStore coerces data stored to bytes.

Looks good.

Just had one further comment re implementation of DictStore.__setitem__().

Make sure that users are only able to add data to the `DictStore`.
Disallow the storing of a nested `DictStore` though.
@alimanfoo
alimanfoo merged commit 8ebb16c into zarr-developers:masterDec 7, 2018
@alimanfoo

Copy link
Copy Markdown
Member

Thanks @jakirkham.

@jakirkham
jakirkham deleted the ensure_DictStore_contains_only_bytes branch December 7, 2018 14:37
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo