[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse. - #7485

Merged
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head
Jan 3, 2025
Merged

[ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.#7485
facebook-github-bot merged 2 commits into
gh/trivedivivek/36/basefrom
gh/trivedivivek/36/head

Conversation

…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jan 3, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/7485

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fdc9997 with merge base 45bb2dd (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
ghstack-source-id: 260053926
Pull Request resolved: #7485
…hing input texel and kernel values for reuse."
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
[ghstack-poisoned]
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D67774359

trviv added a commit that referenced this pull request Jan 3, 2025
…texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
@facebook-github-bot
facebook-github-bot merged commit 3ead5c9 into gh/trivedivivek/36/baseJan 3, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/trivedivivek/36/head branch January 3, 2025 22:47
trviv added a commit that referenced this pull request Jan 4, 2025
…texel and kernel values for reuse. (#7506)
* [ET-VK] Reduced int precision for all int storage in conv pw op to improve performance.
Pull Request resolved: #7447
This diff reduces the precision of all int storage in the conv pw op to improve performance. The code changes include adding the extension GL_EXT_shader_explicit_arithmetic_types_int16 and changing the data type of ints to uint16.
ghstack-source-id: 260166244
@exported-using-ghexport
Differential Revision: [D67674212](https://our.internmc.facebook.com/intern/diff/D67674212/)
* [ET-VK] Minor fix to conv 2d op using wg_size from create_conv2d_global_wg_size to determine local wg size.
Pull Request resolved: #7450
This diff contains changes to the Convolution.cpp file in the Vulkan backend of Executorch. The changes involve updating the code to use the create_conv2d_global_wg_size function to determine the local workgroup size for the convolution operation. This is done to ensure that the correct workgroup size is used for the operation, which can improve performance.
ghstack-source-id: 260166246
@exported-using-ghexport
Differential Revision: [D67676422](https://our.internmc.facebook.com/intern/diff/D67676422/)
* [ET-VK] Modify conv 2d pw op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
Pull Request resolved: #7452
This diff modifies the convolution 2D pointwise op shader and dispatch settings to linearly dispatch work accounting for linearity texture to improve performance.
ghstack-source-id: 260166247
@exported-using-ghexport
Differential Revision: [D67683411](https://our.internmc.facebook.com/intern/diff/D67683411/)
* [ET-VK] Using vec2 to store output positions to reudce shader register footprint.
Pull Request resolved: #7474
The diff changes the use of `u16vec3` to `u16vec2` to store output positions in the conv2d_pw op. This change is made to reduce the shader register footprint and improve performance.
ghstack-source-id: 260166245
@exported-using-ghexport
Differential Revision: [D67726229](https://our.internmc.facebook.com/intern/diff/D67726229/)
* [ET-VK] Using shared variable to store calculated output pose to free up registers and improve performance.
Pull Request resolved: #7475
This diff introduces a shared variable to store calculated output pose in conv2d_pw op to free up registers and improve performance. The code changes include adding a shared variable to hold calculated positions and modifying the existing code to use the shared variable.
ghstack-source-id: 260166242
Differential Revision: [D67742567](https://our.internmc.facebook.com/intern/diff/D67742567/)
* [ET-VK] Changing texture access pattern for conv2d pw op to improve performance.
Pull Request resolved: #7476
This diff changes the texture access pattern for conv2d pw op to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166241
@exported-using-ghexport
Differential Revision: [D67769100](https://our.internmc.facebook.com/intern/diff/D67769100/)
* [ET-VK] Changing texture access pattern for conv2d dw op to improve performance.
Pull Request resolved: #7477
This diff changes the texture access pattern for convolutional depthwise (DW) operations in Executorch's Vulkan backend to iterate first on x axis then y and then z to improve performance.
ghstack-source-id: 260166240
@exported-using-ghexport
Differential Revision: [D67770160](https://our.internmc.facebook.com/intern/diff/D67770160/)
* [ET-VK] Adding batch processing to conv2d dw shader by caching input texel and kernel values for reuse.
Pull Request resolved: #7485
This diff adds batch processing to the conv2d dw shader by caching input texel and kernel values for reuse. This optimization reduces the number of texture lookups and kernel computations, improving the performance of the convolution operation.
ghstack-source-id: 260166243
@exported-using-ghexport
Differential Revision: [D67774359](https://our.internmc.facebook.com/intern/diff/D67774359/)
---------
Co-authored-by: Vivek Trivedi <5340687+trivedivivek@users.noreply.github.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@trviv@facebook-github-bot@SS-JIA