CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

CUDA pipeline for computing APR - #185

Open
krzysg wants to merge 91 commits into
masterfrom
cuda
Open

CUDA pipeline for computing APR#185
krzysg wants to merge 91 commits into
masterfrom
cuda

Conversation

@krzysg

Copy link
Copy Markdown
Member

No description provided.

krzysg added 30 commits August 1, 2022 14:39
…adding), still float number differences between CPU and GPU
@cheesema

Copy link
Copy Markdown
Member

Can't wait to give this a go amazing!

LinearAccess structure on GPU (so it does not support old random or sparse data structures).

Good decision those are all that are needed / suited for the GPU anyhow.

@krzysg
krzysg changed the base branch from develop to masterOctober 30, 2024 09:42
@krzysg

Copy link
Copy Markdown
MemberAuthor

Cool thanks! I think develop is stale, and maybe we should point this at main?

Sure - as we talked lately I have change target branch to 'master'.

@krzysg

Copy link
Copy Markdown
MemberAuthor

Hi, there are two parts of LIS that probably require some explanation:

  1. 'bool boundaryReflect' additional parameter to all calc_sat_mean_*

At some point CPU impl. was change to padd/unpadd pixels before running LIS (according reflect_bc_lis parameter). At first I was trying to avoid it for GPU since I was trying to avoid additional mem allocation for padded pixels. And I managed to do that for 1D. Unfortunately it does not work for 2D or 3D cases (which is obviuos since in CPU impl. cals_sat_mean_* when run also change padded pixels so when you run that in Y-dir it has some influance on running later X-dir and so on).
Anyway I decided to leave it as a option for super fast processing of 1D data (without allocation of additional memory, double mem copy for padd and unpadd). So I decided to upgrade also CPU impl. with that 'feature'.

  1. And now we got second part - as you know I really pay attention (no matter if this is important or not) to correct math and having same results for same data even if some differences would be caused by some minor float operations differences only (like because of different order of execution or so). So in case if I have some date put in Y-dir and when I run calc_sat_mean_y I want to have same results as when I put same data transposed (same date in X-dir) and run calc_sat_mean_x. Unfortunately those functions in CPU impl. where giving little bit different results (+/- some float precision on 6th or 7th digit after decimal point - enough to make me angry ;-) ).
    Since I first managed to have proper behavior on GPU code so I have 'equalized' code on CPU side to GPU side. So the good things because of that are:
    (a) code is still doing same thing that old code (+/- mention float precision) but behaves in a same way in all directions (it is already removed from code but during development I wrote tests comparing outputs of both old/new LIS to check if all is correct).
    (b) naming of variables is same of GPU/CPU side so in case of any changes it is easy to fix/update both sides since they look similar.

krzysg added 26 commits March 17, 2025 16:17
…puted in constructor and memory for them is preallocated
…ixed. Still APRConverter for streams must be fixed/improved since there is a draft code.
…ly one ARP object - use it only for speed for now
…init in APRConverter for GPU not needed anymore
…s and limit log messages for benchmarking purposes
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@krzysg@cheesema