Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); accelerate plotly JSON encoder for numpy arrays without nans by emmanuelle · Pull Request #2880 · plotly/plotly.py · GitHub
Skip to content

accelerate plotly JSON encoder for numpy arrays without nans - #2880

Merged
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder
Nov 17, 2020
Merged

accelerate plotly JSON encoder for numpy arrays without nans#2880
emmanuelle merged 10 commits into
masterfrom
accelerate-encoder

Conversation

@emmanuelle

@emmanuelleemmanuelle commented Nov 6, 2020

Copy link
Copy Markdown
Contributor

This is a proof of concept PR, I need to add a bunch of tests to be sure it behaves well for numpy arrays of different types, but the idea is to remove the encoding-decoding-reencoding process when JSON-serializing a numpy array without nans or infs (which is the reason why there is this 3-step process).

For large data arrays (for example a 1000x1000 array) there is a large performance gain. For example for

fromtimeimporttimeimportnumpyasnpimportplotly.expressaspximportplotly.graph_objectsasgofromplotly.utilsimportPlotlyJSONEncoderimportjsonimg=np.random.randint(255, size=(1000, 1000)).astype(np.uint8)
fig=go.Figure(go.Image(z=img))
t1=time()
_=json.dumps(fig, cls=PlotlyJSONEncoder)
t2=time()
print(t2-t1)

the proposed change results in a 2.5x performance gain on my machine (0.11 s instead of 0.27). The performance gain is interesting in Dash apps in particular.

@jonmmease what do you think of this?

@jonmmease

Copy link
Copy Markdown
Contributor

Is the basic idea here that you're skipping the recursive encoding of the elements of the array? If so, that makes a lot of sense.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

I understand (but maybe this is not correct) that there is an additional cycle of decoding-reencoding here in order to take care of nans and infs. If this is the case, we can check if there are no nans and infs and skip these steps (the "recursive encoding") I think.

But maybe I don't understand completely the purpose of the recursive encoding. The CI failure of one of the builds is definitely linked to this PR, I need to investigate.

@jonmmease

Copy link
Copy Markdown
Contributor

ohh, wow, I didn't realize (remember?) that we did that double encoding. That would definitely slow things down. It would be great to get rid of that all together, but it would take some investigation to work through why it's there in the first place.

What I meant by "recursive" is that JSON serialization handles lists of objects (like fig["data"]) by recursively encoding each element of the list. I'm not certain, but I think this recursive encoding happens for each element of numpy arrays as well, even though it wouldn't be strictly necessary.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

It'd be nice to get rid of the double-encoding in general but in the meantime, if there are special cases where we feel safe to skip it e.g. arrays with every entry a finite value, then that would be a short-term win also for many cases :)

@nicolaskruchten

Copy link
Copy Markdown
Contributor

I'll note that by playing string games with numpy.savetxt I was able to get around a 10% speedup, but it was pretty gross :)

obj.dtype, numpy.number
):
if numpy.all(numpy.isfinite(obj)):
self.unsafe = True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I am missing something. We can test all the numpy arrays that are in the dict being encoded, but we'd not catch nans and infs in any sublists, right?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this will work only for numpy arrays but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed. Or at least this is the intended behaviour :-), maybe I am the one missing something! I need to write tests anyway.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but for other kinds of objects the existing process (encoding - decoding - reencoding) will still be executed

I don't think this is the case, since as I understand this class, encode() is not called recursively, or is it?

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@jonmmease I wonder if you could help me understanding the py3.7-orca CI failure, which I cannot reproduce locally in a conda environment with (I think) the same packages installed and py3.7. The problem occurs in the doctest of the create_dendogram figure factory, there is both an ImportError and a ValueError. Does it make sense to you?

@almarklein

almarklein commented Nov 9, 2020

Copy link
Copy Markdown

FWIW - just in case anyone was looking to go down this route: Iast Friday I did a brief stab at making encoding nan and inf as null in the first place (instead of replacing such values after the fact by the extra decode-encode pass).

I figured we could overload the JSONEncoder.iterencode method. It defines a function floatstr(), which seems exactly the thing we'd want to change. Except ... it's only used when the Python-based low-level encoder is used. It looks like the c-version brings it's own version of floatstr() :(

added: I did a quick test to force using the Python (non-C) version of the builtin encoder. But that makes it twice as slow.

@pfbuxton

pfbuxton commented Nov 13, 2020

Copy link
Copy Markdown

Hi @emmanuelle,

A while ago I profiled plotly/dash and found the numpy to JSON encoding to be one of the bottlenecks and also found that by not doing the encoding twice the code can be sped up by approximately a factor 2:
#1842 (comment)

The way I implemented it I was able to handle numpy arrays with both +/- inf's and nan's without the need to re-encode. But some care is needed as numpy has quite a lot of different variable types.

Hope this helps.

@emmanuelle

Copy link
Copy Markdown
ContributorAuthor

@almarklein you were right of course I had not understood in which order the different methods were called.

My workaround with np.isfinite did not work since some arguments might be something like x=[1, float("nan"), "platypus"] (this is actually a variable used in the tests!) and it's not possible to call np.isfinite on this kind of beast.

Therefore I resorted to a brute-force approach of testing whether the json string encoded the first time has the substrings "NaN" or "Infinity" and if not, skip the additional decoding and re-encoding step. For large data arrays the cost of checking these two patterns if small compared to the encoding time (~2% of the time on a couple of examples I tested). This is a slight cost burden to be added to figures with NaN or Infinity either in their data or in text strings (for example titles, labels etc.) but I think we can live with this, given the large speed-up (x2.5) for non-problematic cases.

@emmanuelleemmanuelle changed the title [WIP] accelerate plotly JSON encoder for numpy arrays without nansaccelerate plotly JSON encoder for numpy arrays without nansNov 15, 2020

@jonmmeasejonmmease left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One question, but otherwise lgtm

"""

# this will raise errors in a normal-expected way
self.hasinfnans = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you using hasinfnans anywhere else now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no you were right :-). Thanks!

# but this is ok since the intention is to skip the decoding / reencoding
# step when it's completely safe
if not ("Infinity" in encoded_o or "NaN" in encoded_o):
return encoded_o

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm fine with this for now given your description of the performance characteristics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe check for NaN first, as it feels like that'd appear more frequently?

@almarklein

almarklein commented Nov 16, 2020

Copy link
Copy Markdown

I like this solution! I find it actually quite elegant given the limitations of the json encoder that we have to work with, and I'm pretty sure that any solution that somehow checks the input beforehand would be slower than this proposed approach.

The only downside is that the speed of encoding depends on the presence of NaN and Inf in the object, which might be an unexpected effect. But these cases are hopefully rare, e.g. the plotly example about gaps in line charts uses None, which becomes null in json.

@nicolaskruchten

Copy link
Copy Markdown
Contributor

💃

@nicolaskruchten

Copy link
Copy Markdown
Contributor

@pfbuxton looks like you went down this path before, I'm sorry no one replied to you more than a year ago :( Thanks for bringing up your idea again in this context!

I think that your idea would layer well onto this one, i.e. now that we only re-encode when we need to, modifying encode_as_list in the way you suggest would safely exploit this fast path.

@emmanuelle
emmanuelle merged commit fa9500b into masterNov 17, 2020
@emmanuelle
emmanuelle deleted the accelerate-encoder branch December 1, 2020 13:45
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@emmanuelle@jonmmease@nicolaskruchten@almarklein@pfbuxton