[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[microNPU][1] Add affine analysis structures for the cascader - #9458

Merged
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1
Dec 1, 2021
Merged

[microNPU][1] Add affine analysis structures for the cascader#9458
leandron merged 4 commits into
apache:mainfrom
mbaret:ethosu-cascader-1

Conversation

@mbaret

Copy link
Copy Markdown
Contributor

RFC: apache/tvm-rfcs#37
Issue: #9429

The cascader relies heavily on being able to determine data dependencies between operators. This is so that it can calculate how stripes should be propagated through a cascade.

To do this, two data structures are defined: StripeConfig and Propagator. StripeConfig stores information for how a tensor should be broken up into stripes and executed. Propagator transforms a StripeConfig using an affine transform matrix, allowing an input StripeConfig for an operator to be determined by 'propagating' the output StripeConfig.

By chaining together Propagators, we can analyse how data dependencies vary throughout a cascade and therefore calculate the memory requirements (and approximate the performance).

@mbaret

Copy link
Copy Markdown
ContributorAuthor

Comment threadsrc/contrib/ethosu/cascader/stripe_config.h
Comment threadsrc/contrib/ethosu/cascader/propagator.cc Outdated
@mbaretmbaret changed the title [ETHOSU][1] Add affine analysis structures for the cascader[microNPU][1] Add affine analysis structures for the cascaderNov 8, 2021
@mbaret

Copy link
Copy Markdown
ContributorAuthor

also cc @csullivan

@tqchen

Copy link
Copy Markdown
Member

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

@mbaret

Copy link
Copy Markdown
ContributorAuthor

minor note: not needed for now, but it may be helpful to take a look at https://github.com/apache/tvm/blob/main/include/tvm/arith/iter_affine_map.h to see if there is any utils that can be reused

I did spend a little bit of time looking at this after @junrushao1994 made me aware of it. It looks like it'd make for a good integration point for a 'v2' - especially once we upgrade to TensorIR. However, I think there are a few representational issues that would make it hard to directly leverage in the current design.

@mbaret
mbaretforce-pushed the ethosu-cascader-1 branch 2 times, most recently from b878cb5 to c71c421CompareNovember 18, 2021 13:40

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great stuff @mbaret, this is only a partial review but I had a couple questions and comments. I'll hopefully be able to wrap the review up on Monday. Apologies for the slow turn around, but thanks for the great contribution!

* error will compound when we need to multiply the strides by the number of
* stripes along a given axis.
*/
inline std::vector<float> GetStrides() const { return strides_; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there an example that we can write or link to in this doc string to illustrate the error accumulation that occurs from using ceildiv style rounding. I didn't quite follow the current description. By fractional striding are you trying to describe the case when the stripe shape dims are not divisors of the input shape dims?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The best example of this is an upscale operation. If we have a 2x2 upscale and choose to stripe the output in 3x3 stripes, the input stripes will get a fractional stride of 3/2. We can't just ceil round this, otherwise as we increase the striding it'll get further and further from the truth. I'll add this to the docs.

Side note: I have considered re-expressing this as a rational number rather than a float, but I think that can be a future improvement for now.

*
* The size of that stripe in each axis is the 'shape'. The strides is how far
* you should move between stripes, so also (4, 4) for a simple non-overlappping
* tiling. However, we explore some overlapping scheduling options so shape != strides

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ping on earlier comment ⬆️ , an example like one of these when stride value is non-integral would be nice.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've added an example based on 2x2 upscale.

*
* Finally, the 'offset' tells us where to start the first stripe. In this simple
* case the offset is just (0, 0), but in something like a padding operation we
* may want to start from a negative index, which is captured by the offset.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note about how negative indexing is handled for an operation would be helpful. For example, the stripe config between two operations Op1 and Op2, will depend on the padding needed by Op2, but will influence where Op1 writes its memory. Is that correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've updated the doc to use slice as the example here because I think that's easier to follow. Regarding the padding case, it's a bit of a challenge to explain without a diagram but I'll give it a go here.

Let's say we have an op A that represents a symmetric pad by 1 and two tensors T_in and T_out such that T_in -> A -> T_out. If T_in has shape (4, 4) then T_out will have shape (6, 6) after the padding. Now, we choose a StripeConfig for T_out which is equivalent to (2, 2) tiling:

StripeConfig T_out = {
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[0, 0],
}

The question then is, what StripeConfig will need to be produced when we propagate to T_in? Well, if we state that any part of a stripe that lies outside of the tensor bounds can be ignored, then the input StripeConfig should be this:

StripeConfig T_in =
{
shape=[2, 2],
extent=[6, 6],
strides=[2, 2],
order=[1, 2],
stripes=[3, 3],
offset=[-1, -1],
}

This will effectively 'overlay' the output StripeConfig on the (4, 4) input tensor T_in, but such that outer 1-wide padding margin is always out-of-bounds (either <0 or >=4 in either axis). This way, reads will never be generated for the padding but they will for the 'interior'.

If this is still unclear - which I appreciate it might be - I can potentially produce a quick diagram.

@manupakmanupak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly data structure introductions at this point! LGTM.

The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
@mbaret

Copy link
Copy Markdown
ContributorAuthor

ping @csullivan

@NicolaLancellottiNicolaLancellotti left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@csullivancsullivan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @mbaret.

@leandron
leandron merged commit 2275359 into apache:mainDec 1, 2021
@leandron

Copy link
Copy Markdown
Contributor

This is merged now, thanks @csullivan@mbaret@NicolaLancellotti@manupa-arm and @tqchen.

masahi pushed a commit to masahi/tvm that referenced this pull request Dec 1, 2021
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 7, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 11, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
yangulei pushed a commit to yangulei/tvm that referenced this pull request Jan 12, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
ylc pushed a commit to ylc/tvm that referenced this pull request Jan 13, 2022
…#9458)
* [ETHOSU][1] Add affine analysis structures for the cascader
The cascader relies heavily on being able to determine
data dependencies between operators. This is so that it
can calculate how stripes should be propagated through a
cascade.
To do this, two data structures are defined: StripeConfig
and Propagator. StripeConfig stores information for how a
tensor should be broken up into stripes and executed.
Propagator transforms a StripeConfig using an affine
transform matrix, allowing an input StripeConfig for an
operator to be determined by 'propagating' the output
StripeConfig.
By chaining together Propagators, we can analyse how
data dependencies vary throughout a cascade and therefore
calculate the memory requirements (and approximate the
performance).
Change-Id: If7176fea961c631be4a6c195303da536030d957b
* Add test guards
Change-Id: I1d7633e20daab33642fa5c4a12e474a4def4d8b8
* Address review comments
Change-Id: Iff5f1effa08e0628de91f5577487d0cecebec824
* Improve docs
Change-Id: I508809d8c1a08d231e3a9b0fd9b3f2639cc2f0e3
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@mbaret@tqchen@leandron@csullivan@NicolaLancellotti@manupak