[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[fb] optimize dithering for WHITE and BLACK - #73

Merged
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither
Feb 1, 2021
Merged

[fb] optimize dithering for WHITE and BLACK#73
raisjn merged 1 commit into
rmkit-dev:masterfrom
mrichards42:optimized-dither

Conversation

@mrichards42

Copy link
Copy Markdown
Collaborator

I've noticed that some of the remarkable_puzzle animations are super slow, so I've been doing some profiling. draw_rect and whichever dithering function I use together take up about 90% of the time according to gprof, so I'm focusing on those :)

One pretty simple optimization is skipping dithering for WHITE and BLACK colors -- the dithering matrices are normalized so that they add or subtract less than 50% of the difference between two colors (so, e.g. in 2-color dithering 0.0 is shifted to between -0.49 and 0.49, which rounds back to 0.0 in all cases).

This change gives me between a 2x and 5x speedup in dithering, depending on how much gray is being drawn:

A sample of lightup, which uses a lot of gray:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 45.86 1.22 1.22 24576951 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
39.47 2.27 1.05 1253836 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
# after
% cumulative self self total time seconds seconds calls ms/call ms/call name 54.08 1.06 1.06 1259815 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
31.12 1.67 0.61 27939892 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
A sample of untangle, which is almost all black and white:
# before
% cumulative self self total time seconds seconds calls ms/call ms/call name 61.84 7.34 7.34 165550334 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)
35.89 11.60 4.26 2498395 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
% cumulative self self total time seconds seconds calls ms/call ms/call name 78.94 5.51 5.51 3671557 0.00 0.00 framebuffer::FB::draw_rect(int, int, int, int, int, int, float)
16.48 6.66 1.15 240161310 0.00 0.00 framebuffer::DITHER::BAYER_2(int, int, unsigned short)

I'd also be happy to add these as separate dithering modes (e.g. BAYER_BW_2 or something), in case we want to let users opt in to this behavior. If you were rendering a grayscale bitmap, for instance, this is probably unhelpful, and just adds 2 extra comparisons per pixel.

These colors should never change due to dithering, and this skips the
costly conversion from remarkable_color to float and back.
@raisjn

Copy link
Copy Markdown
Member

this seems reasonable, thanks!

regarding animations and potential slow down, also see ddvk/remarkable2-framebuffer#42. if the rm2fb queue is filled, it will cause delay in painting until the queue settles (this is contrary to how rm1 worked, i believe).

if you turn on DEBUG in the rm2fb client/server code, you can see the update requests and draw timings - an example output can be seen here: ddvk/remarkable2-framebuffer#38 (comment). this will let you get a sense of how much repainting a particular region costs with a particular waveform.

@raisjn
raisjn merged commit 1ac0cfd into rmkit-dev:masterFeb 1, 2021
@raisjn

Copy link
Copy Markdown
Member

i just played with the release build with these changes, the performance is good! i had been using the debug build previously, so was very happy to see release build perf - i think it feels pretty snappy for an e-reader

(i also think the rm2fb queue getting filled is not the issue after doing some small testing)

@mrichards42

Copy link
Copy Markdown
CollaboratorAuthor

Awesome! Yeah, it's pretty decent with a release build. Agreed that I don't think it's an issue w/ the rm2fb queue -- although that might have been part of the trouble I had w/ the dithering demo.

@raisjn

raisjn commented Feb 3, 2021

Copy link
Copy Markdown
Member

i'm curious if #78 speeds up drawing at all for you

  • uses draw_rect_fast which doesn't call update_dirty (which does a bunch of extra work for each pixel)
  • uses likely/unlikely
  • split rectangle drawing into separate cases: one case for fill, one for non-fill. non-fill is now O(N+M) instead of (N*M)

i am unsure of how you did your profiling cases above - is it just a single render? (or X renders?) of a game scene? or does it involve manual interaction with the game?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@mrichards42@raisjn