Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Arm backend: Int16 linear support - #14258

Merged
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support
Sep 24, 2025
Merged

Arm backend: Int16 linear support#14258
zingo merged 7 commits into
pytorch:mainfrom
per:int16_linear_support

Conversation

@per

@perper commented Sep 12, 2025

Copy link
Copy Markdown
Collaborator

Summary

Adds support for a16w8 for linear when targeting a backend with +int16 extension.

Fixes#13729

Test plan

Tested through unit tests.

cc @digantdesai@freddan80@zingo@oscarandersson8218

@pytorch-bot

pytorch-botBot commented Sep 12, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14258

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit 0bdb35c with merge base 0329a8a (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 12, 2025
@per
per requested a review from zingoSeptember 12, 2025 14:22
@perper added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm ciflow/trunk release notes: none Do not include this in the release notes labels Sep 12, 2025
@zingozingo added this to the 1.0.0 milestone Sep 12, 2025
@zingozingo changed the title Int16 linear supportArm backend: Int16 linear supportSep 12, 2025
@per
perforce-pushed the int16_linear_support branch from e170183 to 0ab797fCompareSeptember 15, 2025 07:08
@per

per commented Sep 15, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Unrelated failures.

Comment threadbackends/arm/_passes/add_bias_pass.py Outdated
self.add_pass(FuseConstantArgsPass(exported_program))
self.add_pass(InsertTableOpsPass(exported_program))
# If we have a conv2d with int16 activation split up into a convolution
# and an addition, to work-around the lack of support for int48 in torch

@digantdesaidigantdesaiSep 15, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and an addition, to work-around the lack of support for int48 in torch

Or can it be done by using torch.dtype.int64 instead and then detecting and lowering it as int48 downstream?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was starting of in that direction, but it interfere a bit with the int64->int32 handling, so rather keep it separate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah given int64 is treated as radioactive :P

@digantdesaidigantdesai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is awesome, No conv tests for 16a8w, just curious.

)


@common.parametrize("test_data", test_data_all_16a8w)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@per - Does it work with U55 and U85?
@Ninja91 - IIRC you already have tests seems like haven't merged?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, not yet. Test will be added for Ethos-U with a Vela update when it's in place.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have added tests for U55 and U85

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo

Copy link
Copy Markdown
Collaborator

I see this testfail :(
[test-arm-backend (test_pytest_ops_ethosu_fvp) / linux-job https://github.com/pytorch/executorch/actions/runs/17768877335/job/50517971935?pr=14258#logs

FAILED backends/arm/test/ops/test_linear.py::test_linear_16a8w_tosa_INT[model_linear_rank1_large_randn,per_channel_quant=True] - AssertionError: Output 0 does not match reference output.
Given atol: 0.004528127912431955, rtol: 0.001.
Output tensor shape: torch.Size([20]), dtype: torch.float32
Difference: max: 0.1481781005859375, abs: 0.1481781005859375, mean abs error: 0.008467483520507812.
-- Model vs. Reference --
Numel: 20, 20
Median: 14.398289680480957, 14.398289680480957
Mean: 2.563008487224579, 2.555246722698212
Max: 69.8498764038086, 69.84635162353516
Min: -115.46151733398438, -115.60969543457031

@per
perforce-pushed the int16_linear_support branch from ac4203c to 54ad8e2CompareSeptember 17, 2025 15:20
@zingozingo removed this from the 1.0.0 milestone Sep 18, 2025
@zingo

Copy link
Copy Markdown
Collaborator

I removed milestone 1.0 for now Vela does not support this anyway so lets take this a bit calmer.

@zingo

zingo commented Sep 19, 2025

Copy link
Copy Markdown
Collaborator

@digantdesai , this has come out of sync with Meta internal version and I'm not allowed merge it, would you like to assist?

@digantdesai

Copy link
Copy Markdown
Contributor

yeah let me try to merge this.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@digantdesai

Copy link
Copy Markdown
Contributor

if it doesn't work I can give you a patch..

Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I2b189b559f699c7eda6921ed515c0e8a849226ca
Support quantization to 16a8w. Since the resulting TOSA operator
needs to have the bias in int48 which isn't avaiable as a type in
torch, the conv2d needs to be decomposed into a conv + add, where
the conv result is scaled down to 32 bit before the addition of the
bias is done.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ib8cae694035796374a55a9909e501596e983abf5
For the case when the activation is 16 bit the bias in TOSA must be
a int48_t tensor. Since that can't be represented using torch.dtypes
the corresponding node.meta is set with a key 'tosa_dtype_48bit' to
pass through the note to the creation of the TOSA Tensor.
Also make sure to distinguish between int32 and int48 tensors in fuse
constant ops pass.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Iefe64f2b02f388c905c9c818ee7d2a6af40bc9e3
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: Ibe158d8d35a632547290f1b9a055d061ae267d77
Enable tests of int16 activations and int8 weight quantization.
Test for large_rand is disabled to sort out why the test is flaky.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I9de5d472f8862edebcf82c140399985db930c069
Add a enum class to handle special dtypes that can't be represented
in torch (i.e. int48_t) to avoid leaking serializer types into the
pass handling of the backend.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Change-Id: I3388cec3c8a26f28790eedc3f124c336b6724cb4
@per
perforce-pushed the int16_linear_support branch from dc7a0a7 to 18c9985CompareSeptember 22, 2025 10:06
@zingo

Copy link
Copy Markdown
Collaborator

Fails are unrelated
Note: Fail in trunk / test-arm-ootb-linux / linux-job is unrelated and handled by another PR

@zingo

Copy link
Copy Markdown
Collaborator

@digantdesai we had to rebase as a file was changed and as we changed it again we got back to the same problem and need your help again. Sorry about that.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@digantdesai has imported this pull request. If you are a Meta employee, you can view this in D82552402.

@zingo
zingo merged commit bb81136 into pytorch:mainSep 24, 2025
359 of 365 checks passed
jirioc pushed a commit to nxp-upstream/executorch that referenced this pull request Dec 19, 2025
### Summary
Adds support for a16w8 for linear when targeting a backend with +int16
extension.
Fixespytorch#13729
### Test plan
Tested through unit tests.
Signed-off-by: Per Åstrand <per.astrand@arm.com>
Co-authored-by: Digant Desai <digantdesai@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: armFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Armrelease notes: noneDo not include this in the release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Arm] Support Linear INT16 TOSA reference model run

5 participants

@per@facebook-github-bot@zingo@digantdesai@Ninja91