[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[XNNPACK] resolve ambiguity around 2d affine quantized tensors - #8958

Merged
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity
Mar 15, 2025
Merged

[XNNPACK] resolve ambiguity around 2d affine quantized tensors#8958
facebook-github-bot merged 1 commit into
pytorch:mainfrom
mcr229:q_affine_ambiguity

Conversation

@mcr229

Copy link
Copy Markdown
Contributor

Summary

There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:

  • input_shape (5, 2048)
  • group_size (1, 2048)

This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.

Test plan

python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4

@mcr229
mcr229 requested a review from kimishpatelMarch 5, 2025 04:30
@mcr229
mcr229 requested a review from digantdesai as a code ownerMarch 5, 2025 04:30
@pytorch-bot

pytorch-botBot commented Mar 5, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/8958

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 89a248a with merge base 1011fdc (image):

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 5, 2025
@mcr229
mcr229force-pushed the q_affine_ambiguity branch from 4febe74 to 0a6b599CompareMarch 5, 2025 04:39
@mcr229
mcr229 requested review from digantdesai and removed request for digantdesaiMarch 5, 2025 04:39

@jackzhxngjackzhxng left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is needed to resolve numerical inaccuracies introduced by switching to the quantize_ API here: #8772. Will leave to @digantdesai / @kimishpatel for final stamp

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@mcr229 has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

Summary:
There is some ambiguity around deriving per_token and per_channel_group semantics from quantize affine. Specifically when rank is two. Take for example the following input shape with the following group size:
- input_shape (5, 2048)
- group_size (1, 2048)
This tensor has 5 scales, and can be seen has having one scale per batch dimension. Or one scale per token. However, this can also be seen as a single weight in which each group is the size of a channel, effectively giving per_channel semantics. This ambiguity does not play well within the XNNPACK Backend as we are must parse these differing quantization types. For now we rely on the fact that per_token quantization happens dynamically. Meaning that the scales and zero points are dynamically choosen. As a result, we check that the scales come from getitem and is dynamically chosen. We further ensure that per_channel_group checks are not per_token.
Test Plan:
```
python -m unittest backends.xnnpack.test.ops.test_linear.TestLinear.test_linear_qd8_f32_per_token_weight_per_channel_group_int4
```
Reviewed By: GregoryComer
Differential Revision: D70719546
Pulled By: mcr229
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70719546

@facebook-github-bot
facebook-github-bot merged commit 0ccf509 into pytorch:mainMar 15, 2025
DannyYuyang-quic pushed a commit to CodeLinaro/executorch that referenced this pull request Apr 2, 2025
Differential Revision: D70719546
Pull Request resolved: pytorch#8958
@mcr229
mcr229 deleted the q_affine_ambiguity branch July 25, 2025 22:43
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedtopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@mcr229@facebook-github-bot@GregoryComer@jackzhxng