Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Coerce data to text for JSON parsing - #429

Merged
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type
Apr 19, 2019
Merged

Coerce data to text for JSON parsing#429
jakirkham merged 11 commits into
zarr-developers:masterfrom
jakirkham:add_ensure_text_type

Conversation

@jakirkham

@jakirkhamjakirkham commented Apr 15, 2019

Copy link
Copy Markdown
Member

Cleans up some Python 2/3 code for handling JSON parsing by simply always coercing metadata to text regardless of Python version.

xref: #372
xref: #401

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

To simplify the branching required for Python 2/3 compatibility.
Rename `ensure_str` to `ensure_text_type` and rework the code to
coerce data that is `bytes` or `bytes`-like to `bytes` and then to
text data. It appears JSON on Python 2 or Python 3 handles this just
fine. So should make handling these two cases a bit more
straightforward.
`MongoDBStore` inherited the behavior on `pymongo` with respect to
returning `bson.Binary` for blob values on Python 2. As this caused
some issues on Python 2 when parsing JSON content (as the parser was
unable) to work with objects that were not `bytes` type (i.e.
`bson.Binary`), a workaround was needed to coerce `bson.Binary` to
`bytes` on Python 2. It's worth noting that this workaround is not
needed for loading binary data from chunks as we use the buffer
protocol there.
As we have now fixed our handling of JSON data to coerce data to text
on Python 2/3 and leverage the buffer protocol in the effort, we no
longer need this workaround in `MongoDBStore`. Hence we go ahead and
drop it.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Comment threadzarr/storage.py
value = binary_type(value)

return value
return doc[self._value]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes me very happy to see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Likewise. 🙂

FWIW it turns out this is not Python 2 specific. We just only handled decoding before parsing JSON on Python 3 (hence avoiding the issue there). With this change we just always decode to text before parsing JSON. Here's a short reproducer.

>>>importjson>>>json.loads(b"{}")
{}
>>>json.loads(b"{\x00}")
Traceback (mostrecentcalllast):
File"<stdin>", line1, in<module>File"/Users/jkirkham/miniconda/lib/python3.7/json/__init__.py", line348, inloadsreturn_default_decoder.decode(s)
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line337, indecodeobj, end=self.raw_decode(s, idx=_w(s, 0).end())
File"/Users/jkirkham/miniconda/lib/python3.7/json/decoder.py", line353, inraw_decodeobj, end=self.scan_once(s, idx)
json.decoder.JSONDecodeError: Expectingpropertynameenclosedindoublequotes: line1column2 (char1)

Much as we have a helper function for writing JSON, this adds a
helper function for loading JSON. Mainly it ensure data is coerced to
text before handing it off to the JSON parser. Should simplify code
that is loading JSON.
Changes other library code to use `json_loads` for handling text
encoding and JSON parsing. Should simplify things a bit and avoid
having some errors sneak in.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging end of day Friday if no comments.

@meggart

Copy link
Copy Markdown
Member

Maybe this is completely unrelated, but since you touched the JSON code anyway, is there a chance that #412 gets fixed during the process?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

It’s unrelated. Though I agree it’s important to fix. Let’s discuss after.

@jakirkham
jakirkham merged commit a7546b7 into zarr-developers:masterApr 19, 2019
@jakirkham
jakirkham deleted the add_ensure_text_type branch April 19, 2019 23:07
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks all 😄

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@meggart@jhamman