Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler - #37

Merged
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc
Jan 25, 2022
Merged

Introduce the Arm(R) Ethos(TM)-U Cascading Scheduler#37
areusch merged 3 commits into
apache:mainfrom
mbaret:cascading-rfc

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

No description provided.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
@junrushao

Copy link
Copy Markdown
Member

Thanks for the RFC!

  • The TE scheduling part (rolling-buffer) looks good to me.
  • Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you
  • We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated
Comment threadrfcs/0037-arm-ethosu-cascading-planner.md Outdated

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

  1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
  2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

$$stripe_{in} = {M} \cdot {stripe_{out}}$$
![meta-schedule-workflow](../resources/cascading-formula-1.png)

Let's briefly consider how to derive such a transform matrix for a 3x3 unstrided, undilated and unpadded NHWC convolution. Immediately, the '3x3' kernel tells us something important: a single element in the output depends on 3x3 elements in the height/width of the input. If we were instead to consider a 2x2 region of the output in the height/width dimensions, we'd then need a 4x4 region in the input. So in general, the rule is that we need 2 more elements in height and width when calculating the dependencies of an output stripe. It can be shown that more generally this number is the kernel_size-1 in each axis. Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel. This is because the input channel axis is a reduction axis in a convolution, in a sense it isn't 'reflected' in the output. Combining these two observations, we arrive at the following transform matrix:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now to consider the channels, in a convolution no matter how many output elements you are computing you'll always need every input channel.

just curious: what about depthwise convolutions?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In a depthwise you're right that this will be different, but depthwise will also use a different transform matrix. So here we're just working out the transform matrix for a standard 2D convolution.

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Thanks for taking a look @junrushao1994!

Grouping operators together might be challenging in Relay fusion and lowering, and would love to hear more thoughts from you

For Ethos-U, we have a different compilation flow to 'standard' TVM. In particular, those operators which are being offloaded onto Ethos-U don't go through FuseOps but instead are lowered into TE as one large graph. This is what allows us to avoid the problems of Relay level fusion. I agree that were we to want to generalize this technique further, we would need a solution for this.

We have some basic affine analysis infra (iter-affine-map) in TVM, and it would be great if the infra could be reused and improved upon the usecases in your RFC

I hadn't actually noticed the affine analysis infra before, so thanks for pointing me towards this. From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR. Perhaps this could be better discussed on a draft PR to demonstrate specifically what we do with the affine transforms and whether the existing infra would be able to achieve the same effect.

@junrushao

Copy link
Copy Markdown
Member

Thank you for your response @mbaret 😄

From my initial look over it, I suspect there may be elements we could reuse, but the representations are unfamiliar to me as it how to work with them from outside TIR

For sure, we are happy to work together and potentially improve the affine analysis infrastructure. Note that the iter-affine-map is located under tvm/arith/iter_affine_map.h, so it's not deeply coupled with TIR, but only reuses several TIR expressions.

CC'ed the author of affine analysis in TVM: @tqchen@spectrometerHBH@Hzfengsy

Change-Id: I34d0acf303391157cf3ddab9ae771b3f6c667967
Change-Id: Ie4b6fbbc0cc060da6a733c06f276f2c22d2cbc93
Change-Id: I89e9128a7e18791671fb4a7baeee9b64e65174c4
@mbaret

Copy link
Copy Markdown
ContributorAuthor

@mbaret thanks for the detailed RFC! i think my main questions here are whether or not this could be implemented in roughly two pieces similar to USMP:

1. a piece which could be reused if we e.g. wanted to attempt to apply this to non-Ethos-U kernels
2. the Ethos-U-specific logic.

I understand it may not make sense to integrate with AutoTVM--it might make more sense to wait til TensorIR has fully landed. but just curious to understand in planning terms what the lift might be.

Apologies for the late reply @areusch. Due to some of the uncertainties around implementing this generically (Relay fusion/conflict with AutoTVM/deriving matrices), I think it's more pragmatic to have everything initially scoped just to Ethos-U. But within that implementation, all the truly Ethos-U specific logic is quite self-contained (pretty much entirely in the performance heuristics and matrix definitions). Therefore, should this need reusing in future we could probably lift it out of the Ethos-U namespace. I think this would be very doable as the cascader itself is quite well decoupled from the rest of the Ethos-U compiler, exposing only a TE Schedule -> TE Schedule interface. With some effort, although hopefully not too much, that interface could likely be tweaked to be TensorIR -> TensorIR.

@mbaretmbaret changed the title Introduce the Arm(R) Ethos(TM)-U Cascading PlannerIntroduce the Arm(R) Ethos(TM)-U Cascading SchedulerNov 4, 2021

@mbs-octomlmbs-octoml left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Mike, thx for such a nicely written RFC, you're upping the bar.

I think we all agreed in our last live conversation:

  • Though the S-TIR machinery is now in main and could be a viable medium instead of TE that transition will be left to future work.
  • The transition to TE should hopefully not require duplicating the ScheduleBuilder visitor -- I think someone else at ARM was getting that in place. Sorry I can't keep up.
  • I'm hoping we can make constants first class so there's no need for any side inputs. Eg just look for them bound as globals in the IRModule. But that's a refactor which will probably come after this work.
  • There's an ongoing tension between working in Relay+TIR vs TIR-only, but the work-exclusively-in-TE/TIR approach is consistent with both the current AOT flow and the USMP analysis work.

If I got that right perhaps capture it in the 'alternates considered' or someplace.

I can't say anything about possible reuse of the affine x-form machinery when doing the abstract interpretation for the memory footprint. Sounds like it might be similar to the S-TIR situation: maybe not now but can go back later.

So LGTM from me.

@areuschareusch left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks @mbaret we should merge this now :)

@areusch
areusch merged commit f9fa824 into apache:mainJan 25, 2022
@tqchen

tqchen commented Jan 27, 2022

Copy link
Copy Markdown
Member

To followup on this RFC, @mbaret would be great to start followup with a rolling-buffer or related primitive in TensorIR. So we can smoothly transition the solution as we start to migrate to TIR schedule

@mbaret

Copy link
Copy Markdown
ContributorAuthor

We don't have any plans in the short term to look at extending the rolling buffer primitive into TensorIR, but I'd be happy to provide support/assistance to anyone interested in implementing it. I would hope the effort would not be too great as it's already implemented as a TIR pass.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@junrushao@tqchen@areusch@NicolaLancellotti@mbs-octoml