Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Support more ops and verify more models - #14106

Merged
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops
Sep 11, 2025
Merged

Support more ops and verify more models#14106
SS-JIA merged 17 commits into
pytorch:mainfrom
Jiseong-oh:expands_more_ops

Conversation

@Jiseong-oh

@Jiseong-ohJiseong-oh commented Sep 9, 2025

Copy link
Copy Markdown
Collaborator

Expands more ops and Verify more models

Summary

  • Added for more OPs currently supported by Exynos
  • Added functionality for models for Showcase
  • Additional Module for Optimization
  • Fixed ci issue

Test plan

models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit, mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model -c E9955

cc @SS-JIA@digantdesai@kimishpatel

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
@pytorch-bot

pytorch-botBot commented Sep 9, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/14106

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 37 Pending, 2 Unrelated Failures

As of commit 3437e4c with merge base a89b858 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 9, 2025
@Jiseong-ohJiseong-oh added partner: samsung For backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsung release notes: exynos labels Sep 9, 2025
Jiseong-ohand others added 9 commits September 10, 2025 07:01
- Add ops in op builders
- batchmatmul, div, maximun, minimun, rsqrt, slice_copy, sqrt, to_copy
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
- add squeeze, sub ops
Co-authored-by: chong-chen <chong.chen@samsung.com>
1. Propagate constant and remove constant ops.
2. Prevent some ops from decomposing.(HardSwish, linear in mv3)
3. Add models name to support list (ic4, edsr enabled at the same time) in aot_compiler
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Add prelu, softmax, layer_norm, upsample builders.
Add these ops to list which keep ops not to decompose
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
1. Add necessary ops - gelu, expand, etc.
2. Preprocess mul/add/sub scalar ops, add ReplaceScalarOps pass.
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Convert conv1d to conv2d and support logsoftmax op.
w2l can be enabled in samsung enn backend
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Support the ops of bert float model (finetune)
Co-authored-by: chong-chen <chong.chen@samsung.com>
Implement test for each op supported in builders
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
add more model tests to test workflow. Like mv3, dl3, vit, etc.
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Comment threadexamples/samsung/aot_compiler.py Outdated
model = model.eval()
outputs = model(*example_inputs)

print("start start ...")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's removed.

Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
# Output from pretrained model exceeds the representation scale of half-float.
# Finetune bert model on specific task and make output more reasonable for hardware.
# Here is an example.
class MobileBertFinetune:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very cool/useful to have this script as an example, by the way

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Please remove the print statement and I will merge.

@SS-JIA
SS-JIA merged commit 63481e3 into pytorch:mainSep 11, 2025
271 of 277 checks passed
Comment on lines +24 to +25
def __init__(self, *args) -> None:
super().__init__(*args)

@swolchokswolchokSep 11, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think __init__ definitions like this can just be omitted

Comment on lines +30 to +40
input1 = node.args[0]
input_id_1 = self.define_tensor(input1, enn_graph, vals_to_ids)
input2 = node.args[1]
input_id_2 = self.define_tensor(input2, enn_graph, vals_to_ids)

# output
output_id = self.define_tensor(node, enn_graph, vals_to_ids)

enn_graph.define_op(
node.name, "BATCH_MATMUL", [input_id_1, input_id_2], [output_id]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like there's a lot of boilerplate in these visitor definitions. You could package up a few helper subclasses like UnaryOpVisitor, BinaryOpVisitor, etc. that get the operator name ("BATCH_MATMUL" etc.) from a class property similar to the existing target property, and then also accommodate the ones with params by having the helper subclass call self.get_params() (default implementation that returns None on the helper subclass) and pass the result to define_op if it isn't None.

Comment on lines +19 to +25
class LogSoftmax(torch.nn.Module):
def __init__(self, dim) -> None:
super().__init__()
self.module = torch.nn.LogSoftmax(dim=dim)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.module(x)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this looks like it could be deleted and replaced with using torch.nn.LogSoftmax directly

StrycekSimon pushed a commit to nxp-upstream/executorch that referenced this pull request Sep 23, 2025
Expands more ops and Verify more models
### Summary
- Added for more OPs currently supported by Exynos
- Added functionality for models for Showcase
- Additional Module for Optimization
- Fixed ci issue
### Test plan
models="ic3, ic4, resnet18, resnet50, mv3, edsr, dl3, w2l, vit,
mobilebert
python -m executorch.examples.samsung.aot_compiler --model_name=$model
-c E9955
cc @SS-JIA@digantdesai@kimishpatel
---------
Signed-off-by: jiseong.oh <jiseong.oh@samsung.com>
Co-authored-by: chong-chen <chong.chen@samsung.com>
Co-authored-by: jingya-zhang <jingya.zhang@samsung.com>
Co-authored-by: xz-linghu <xz.linghu@samsung.com>
Co-authored-by: Jonghun Cha <jhbb.cha@samsung.com>
@Jiseong-oh
Jiseong-oh deleted the expands_more_ops branch October 10, 2025 04:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.partner: samsungFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Samsungrelease notes: exynos

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Jiseong-oh@swolchok@SS-JIA