Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Arm backend: Remove fast scratch part for now - #10958

Merged
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323
May 19, 2025
Merged

Arm backend: Remove fast scratch part for now#10958
facebook-github-bot merged 2 commits into
mainfrom
export-D74939323

Conversation

@kirklandsign

Copy link
Copy Markdown
Contributor

Summary: Fix CI

Differential Revision: D74939323

Summary: Fix CI
Differential Revision: D74939323
@pytorch-bot

pytorch-botBot commented May 17, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/10958

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 116eaf2 with merge base 9aaea31 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 17, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D74939323

@zingo

Copy link
Copy Markdown
Collaborator

Hi, sorry for the break, from our (Arm) point just do what is easier for you to just make it work. Revert or this fix and we will look into this when back to work next week.

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

@zingozingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, to just make it work for you.

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

To understand the problem you see better and if you know, is it so that you have a different runner then our example runner you are testing with?

Seems that our internal CI has a different runner than the example runner.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@kirklandsign has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@zingo

Copy link
Copy Markdown
Collaborator

** > Seems that our internal CI has a different runner than the example runner.**

Thank that makes sense as you get this problem, and something we will try to keep in mind.

@zingo

zingo commented May 18, 2025

Copy link
Copy Markdown
Collaborator

Regarding the failed Arm unit model test it is unrelated an a fix for it is (hopefully) here
#10953
Just waiting for review

@zingozingo added ciflow/trunk module: arm Issues related to arm backend labels May 18, 2025
@zingozingo changed the title Remove fast scratch part for nowArm backend: Remove fast scratch part for nowMay 18, 2025
@facebook-github-bot
facebook-github-bot deleted the export-D74939323 branch May 19, 2025 00:33
@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

Hi @zingo seems that the test-arm-backend (test_models_ethos-u85) / linux-job always failed after this patch.

https://github.com/pytorch/executorch/actions/runs/15101696584/job/42443518087

Could you please suggest how to fix?

@gggekov

Copy link
Copy Markdown
Collaborator

Hi @kirklandsign,
Yes, when you pass nullptr & 0 as base address/base address size for the U85 in Dedicated_Sram memory mode, the inference hangs as you see in your CI trace - that behaviour is expected. Dedicated_Sram means that the NPU uses the fast scratch buffer as a buffer to store the most commonly accessed intermediate tensors. With your fix, you are still testing Dedicated_Sram on U85, but you are not providing fast scratch array, hence the NPU hangs.

Can you please expand a bit more on what you mean when you say you use a different arm_executor_runner.cpp in your internal test suite - do you have a way to define the ethosu_fast_scratch/ethosu_fast_scratch_size to be 0 or 384KB depending on if we test U55(Shared_Sram) or U85(Dedicated_Sram) in your internal arm_executor_runner.cpp?

The U85 is designed to enable NNs where the scratch buffer is too big to fit into the SRAM and hence we place it in the DDR. In order to achieve good performance, it is important to still "dedicate" small amount of SRAM where the NPU can R/W the most commonly accessed tensors of the NN. That behaviour is enabled by the Dedicated_Sram memory mode, so it's important we support it properly.

gggekov added a commit to gggekov/executorch that referenced this pull request May 19, 2025
Temporary solution to the problem in pytorch#10958
The arm_executor_runner.cpp need to declare the ethosu_fast_scratch array and
pass it onto to the EthosUBackend.cpp. It is important that for Shared_Sram,
the ethosu_fast_scratch is nullptr and for Dedicated_Sram it points to the
fast memory array.
Change-Id: I808203fb7b9b6e5bece92c4cc5079f22bd802d95
@hsharma35

Copy link
Copy Markdown
Contributor

@kirklandsign@gggekov Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@kirklandsign

Copy link
Copy Markdown
ContributorAuthor

@zingo@hsharma35 Sorry I am not the best contact for this issue. I don't work on this and just ran into an internal error. Could use @digantdesai 's help

zingo pushed a commit that referenced this pull request May 19, 2025
…Ethos-U85 (#10973)
Temporary solution to the problem in
#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
Differential Revision: D74939323
Pull Request resolved: pytorch#10958
hinriksnaer pushed a commit to hinriksnaer/executorch that referenced this pull request May 19, 2025
…Ethos-U85 (pytorch#10973)
Temporary solution to the problem in
pytorch#10958 The
arm_executor_runner.cpp need to declare the ethosu_fast_scratch array
and pass it onto to the EthosUBackend.cpp. It is important that for
Shared_Sram, the ethosu_fast_scratch is nullptr and for Dedicated_Sram
it points to the fast memory array.
@gggekov

gggekov commented May 21, 2025

Copy link
Copy Markdown
Collaborator

Hi @hsharma35,

The EthosUBackend.cpp currently doesn't know if the NN has been compiled for Shared_Sram or Dedicated_Sram or another memory mode, but the arm_executor_runner.cpp knows that thanks to the propagation of parameters in the executorch/examples/arm/executor_runner/CMakeLists.txt. For pte generated for Shared_Sram & Sram_Only, we only need to pass 2 base pointers/base pointer sizes, but for Dedicated_Sram we need to pass a third base pointer towards the fast scratch array, hence I declared the fast scratch in the arm_executor_runner depending on the memory mode. Then, i passed the fast scratch(nullptr or valid address) as an extern at link time to the EthosUBackend.cpp

What's the best way to enable dedicated_sram in the runtime & make your internal CI pass ?

CC @digantdesai@kirklandsign

@digantdesai

Copy link
Copy Markdown
Contributor

Can these variable be added to compile spec instead or declaring them in arm_executor_runner.cpp?

@hsharma35 not sure if compile_spec is the way to do this, esp if the delegate runtime wants to know the pointer at runtime based on the runner setup.

We have a flag for specifying the memory_mode from PTE generation time, same used by the CMake IIUC. But for the scratch ptr, size info originates at runtime in the runner (through user specified cmake knobs), and we can't easily pass it in the delegate::init() from a runner.

Here @gggekov used externs but I am not sure if that's the right approach because it leads to precisely these issues where you are now connecting a runner with a delegate - while they shouldn't know about each other given they sit in a different abstraction layers.

We are working on et::backend_configs which may help in this but it is not ready yet. #10216

Let me think some more. For now can we guard the delegate runtime variables such that it is backwards compatible for some runner which doesn't care about this memory_mode? @gggekov

@gggekov

Copy link
Copy Markdown
Collaborator

Thanks @digantdesai , you are right the extern couples the arm_executor_runner.cpp & EthosUBackend.cpp in a way that we should avoid.

How about if we use a weak symbol for the fast scratch array, by default set to it nullptr in the EthosUBackend.cpp? That corresponds to Shared_Sram for U55,65 & U85. In case of a U85 SoC where we want to use the DRAM(Dedicated_Sram), overwrite the weak symbol in the arm_executor_runner.cpp with the correct array for the cache. I believe this should pass your internal CI and is close to a real system where for Dedicated_Sram, the user needs to carve out specified amount of memory for the NPU.

Regarding #10216, in principle it may be useful, but I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85. Also, note that in executorch/backends/arm/arm_vela.py, in the vela npz file you have the size of the scratch buffer and the fast scratch buffer. For Shared_Sram/Sram_Only, the fast_scratch_buffer is 0 and for Dedicated_Sram, it is equal to the amount of SRAM that the NPU can utilise.

@digantdesai

Copy link
Copy Markdown
Contributor

I think weak symbol override is fine. As long as we don't tightly couple the delegate runtime vs. app runner.

I don't think we need a backend specific configuration to enable the Dedicated_Sram mode on the U85

Generally speaking, I see this mechanism as a cleaner mechanism to pass information from outside to a delegate especially when it originates at runtime. For compile-time information we can use other mechanisms like compile_spec, preprocessor macros/variables, etc.

@gggekov

Copy link
Copy Markdown
Collaborator

FYI - here is a pr(#11459) with the suggested fix with the weak symbol

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunkCLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedmodule: armIssues related to arm backendrelease notes: noneDo not include this in the release notestopic: not user facing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@kirklandsign@facebook-github-bot@zingo@gggekov@hsharma35@digantdesai