Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Support extra ops modes for LLM Models - #18670

Merged
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes
Sep 2, 2026
Merged

Support extra ops modes for LLM Models#18670
Jiseong-oh merged 4 commits into
mainfrom
extra_ops_modes

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Apr 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Extra Ops for Exynos Backend

  • Cos, GroupNorm, Index, Log, Pow, Rms_Norm, Sigmoid, sing, Splite_with_sizes_copy, sum_int_list, tanh and Topk
  • All ops are verfied
  • These ops will be used for supporting LLM models

Support Perf mode

  • It can be supported Perf mode with experimental
  • This mode MUST BE used to verify model on exynos device farm fristly before testing on the phone

cc @SS-JIA@digantdesai@kimishpatel

@pytorch-bot

pytorch-botBot commented Apr 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/18670

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 5 Pending

As of commit 2694d01 with merge base 4f689f3 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Apr 2, 2026
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung module: samsung CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. and removed CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. labels Apr 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@SS-JIA

Copy link
Copy Markdown
Contributor

@claude help review this PR

@claude

claudeBot commented Apr 9, 2026

Copy link
Copy Markdown

Claude finished @SS-JIA's task in 4m 41s —— View job


Review of PR #18670: Support extra ops modes for LLM Models

  • Gather context and understand the PR changes
  • Review new op builders
  • Review compose_rms_norm pass
  • Review test files
  • Review serialization and compile options changes
  • Review example script changes and perf mode support
  • Post comprehensive review

Overview

This PR adds 12 new op builders for the Samsung Exynos backend (Cos, GroupNorm, Index, Log, Pow, RmsNorm, Sigmoid, Sin, SplitWithSizesCopy, SumIntList, Tanh, TopK), a RecomposeRmsNorm pass, and performance mode support via a new PerformanceMode enum. Each new op has a corresponding unit test. All example scripts are updated to use PerformanceMode.HIGH_PERFORMANCE by default.

Overall the PR is well-structured and follows existing patterns. Below are the issues I found, organized by severity.


Issues

1. Bug: op_pow.py target vs test mismatch — backends/samsung/builders/op_pow.py:20

The builder registers target = "aten.pow.Tensor_Tensor", but the test in test_pow.py:36 checks for torch.ops.aten.pow.Tensor_Scalar (using x**2.0 which produces Tensor_Scalar). Additionally, the partitioner adds exir_ops.edge.aten.pow.Tensor_Scalar to SUPPORTED_OPS (enn_partitioner.py:42).

This means the builder only handles Tensor_Tensor pow, but the test and partitioner expect Tensor_Scalar. The Tensor_Scalar variant would match via the partitioner's SUPPORTED_OPS list and presumably be handled by the ReplaceOpsWithScalar pass, but the builder itself would never directly match Tensor_Scalar ops. If both variants need builder support, the builder's target should be a list including both. If only Tensor_Tensor is intended for the builder and Tensor_Scalar is handled through the scalar replacement pass, the test should be updated to test the Tensor_Tensor case that actually exercises the builder.

Fix this →

2. compose_rms_norm.py — potential _get_eps_node / _get_gamma_node returning Nonebackends/samsung/_passes/compose_rms_norm.py:22-33

def_get_eps_node(self, nodes):
add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]
forainadd_node.args:
ifisinstance(a, float) ora.op!="call_function":
returnadef_get_gamma_node(self, output_node):
forainoutput_node.args:
ifa.op!="call_function":
returna

If no arg matches the condition, these methods implicitly return None, which would then be passed as the eps/gamma argument to rms_norm. Consider raising an error if the expected node is not found rather than silently passing None.

Fix this →

3. compose_rms_norm.py — fragile node name matching — backends/samsung/_passes/compose_rms_norm.py:24

add_node= [nforninnodesifhasattr(n, "name") and"add"inn.name][0]

Matching by "add" in n.name is fragile — node names like "add_padding", "broadcast_add", etc. could be false positives. The Qualcomm version (backends/qualcomm/_passes/recompose_rms_norm.py) matches by checking node.target against known ATen add ops, which is more robust. Consider matching on node.target instead of node.name.

Fix this →

4. Hardcoded high-performance mode in all examples — examples/samsung/aot_compiler.py:81

Every example script now unconditionally uses PerformanceMode.HIGH_PERFORMANCE. The PR description warns "This mode MUST BE used to verify model on exynos device farm firstly before testing on the phone." But none of the scripts expose this as a CLI argument, so users can't opt out. The aot_compiler.py (the general-purpose compiler) especially should probably default to DEFAULT or expose --perf_mode as an argument.

5. Copyright header — backends/samsung/_passes/compose_rms_norm.py:1

The file has # Copyright (c) Qualcomm Innovation Center, Inc as the first copyright line. This file is derived from backends/qualcomm/_passes/recompose_rms_norm.py (same class name, similar approach), using a get_source_partitions-based implementation. The Qualcomm copyright is appropriate if code was derived from their work, but verify this dual-copyright header is intentional.

6. gen_samsung_backend_compile_spec type hint — backends/samsung/serialization/compile_options.py:76

defgen_samsung_backend_compile_spec(
chipset: str,
perf_mode: PerformanceMode=None,
):

The type hint says PerformanceMode but the default is None. Should be Optional[PerformanceMode] = None for correctness.

7. Inconsistent op name casing in define_op calls

Most existing op builders use ALL_CAPS for the op type string (e.g., "SQRT", "SIGMOID", "TANH", "LOG"), but some new ops use mixed casing:

  • op_sin.py:31: "Sin" (vs "SIN")
  • op_cos.py:31: "Cos" (vs "COS")
  • op_topk.py:74: "TopK" (vs "TOPK")

The existing op_hardsigmoid.py also uses "HardSigmoid" (mixed case), so this may be intentional per the ENN backend's expected op names. Worth confirming these are the exact strings the backend library expects.

8. Minor: op_split_with_sizes_copy.py missing return type annotation — line 26

The define_node method is missing the -> None return type annotation, unlike all other new builders.

9. Minor: op_topk.py shadows built-in sorted — line 70

sorted=cast(bool, node.args[4])

Shadows the Python built-in sorted. Not a functional bug but not great practice.


Positive observations

  • Tests cover multiple configurations (e.g., index on different axes, split with different chunk sizes, sum with/without keepdims, topk with different dims).
  • The RecomposeRmsNorm pass correctly uses get_source_partitions with torch.nn.RMSNorm source matching, which is a clean approach.
  • The FlatBuffers schema change for PerformanceMode is clean and backward-compatible (default = 0).
  • The @experimental decorator on PerformanceMode is a good way to signal that this feature is not yet stable.

Summary

The main issues to address are:

  1. The pow target mismatch (issue Add support for quantized LeakyReLU #1) — verify whether the builder covers the right variant and ensure the test exercises the builder path
  2. Fragile RMS norm pass (issues Re-sync with internal repository #2, Rename _pt2e to pt2e #3) — _get_eps_node/_get_gamma_node could return None and name-based matching is fragile
  3. Hardcoded high-performance mode in all examples (issue Add unlifting pass under private config #4) — consider making this configurable or defaulting to DEFAULT
  4. Type hint correctness (issue Re-sync with internal repository #6) — Optional[PerformanceMode]

The rest are minor style/consistency observations.


@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp to unblock -- but the agent review comments are valid. Would highly recommend fixing them before landing!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 6, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@github-actions

Copy link
Copy Markdown

Looks like this PR hasn't been updated in a while so we're going to go ahead and mark this as Stale.
Feel free to remove the Stale label if you feel this was a mistake.
If you are unable to remove the Stale label please contact a maintainer in order to do so.
If you want the bot to never mark this PR stale again, add the no-stale label.
Stale pull requests will automatically be closed after 30 days of inactivity.

@github-actionsgithub-actionsBot added the Stale PRs inactive for over 60 days label Aug 6, 2026
@Jiseong-ohJiseong-oh removed the Stale PRs inactive for over 60 days label Aug 31, 2026
@Jiseong-oh

Copy link
Copy Markdown
CollaboratorAuthor

/easycla

Comment threadbackends/samsung/builders/op_topk.py
Comment threadbackends/samsung/builders/op_split_with_sizes_copy.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadbackends/samsung/_passes/compose_rms_norm.py
Comment threadbackends/samsung/builders/op_group_norm.py
Comment threadexamples/samsung/scripts/resnet18.py
Jiseong-ohand others added 4 commits September 2, 2026 09:55
The previously supported single ops are being migrated to the current
dev branch, as follows.
cos, group_norm, index, log, pow, rms_norm, sigmoid, sin,
split_with_sizes_copy, sum_int_list, tanh and topk.
Co-authored-by: xz.linghu <xz.linghu@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
LiteCore provide a new api named "graphgen_set_perf_mode".
This commit invokes this api to set performance mode.
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Internal IR remove "input_type" from Gather op.
This commit removes "input_type" setting and set Gather inputs
Co-authored-by: Jingya Zhang <jingya.zhang@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
- split test into PowScalar and PowTensor
- raise RuntimeError when expected nodes aren't found.
- fix perf mode on aot_compiler.py.
- added -> None to define_node.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@Jiseong-oh
Jiseong-oh merged commit 80a07e6 into mainSep 2, 2026
219 checks passed
@Jiseong-oh
Jiseong-oh deleted the extra_ops_modes branch September 2, 2026 11:07
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: samsungpartner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Jiseong-oh@SS-JIA