[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu - #16106

Merged
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule
Nov 24, 2023
Merged

[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpu#16106
lhutton1 merged 6 commits into
apache:mainfrom
Anndrey24:fp32-hybrid-schedule

Conversation

@Anndrey24

@Anndrey24Anndrey24 commented Nov 10, 2023

Copy link
Copy Markdown
Contributor

Implemented an arm_cpu conv2d NHWC schedule using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
In the fp32 case, the micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.

There are now two ways to transform the weights matrix for conv2d, which are detailed in convolution.cc:

  • for int8 / uint8: tile, interleave, transpose
  • for everything else: tile, interleave

To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of tile_rows_B or tile_cols_B have been changed to tile_N and tile_K respectively to denote the tiling size along each axis of the flattened B matrix. As usual, N = out_channels and K = kernel_width * kernel_height * in_channels.

I have also added a new conv2d NHWC fp32 test for both the conv2d_nhwc_spatial_pack and conv2d_NHWC_hybrid schedules, as well as new fp32 and fp16 implementation selection tests in test_select_implementation.py.

cc @ekalda@lhutton1@neildhickey@leandron

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24, this looks like a great change :) I had a couple of comments around selecting when to use the schedule to ensure we don't break existing functionality. I think it would be good to check this by adding some tests to test_select_implementation.py to make sure we're still selecting the correct schedule for different devices.

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/topi/arm_cpu/arm_utils.py Outdated
Comment threadpython/tvm/topi/arm_cpu/conv2d_gemm.py
Comment threadtests/python/integration/test_arm_aprofile.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@Anndrey24Anndrey24 changed the title [TOPI][Relay] Add conv2d NHWC fp32 hybrid schedule for arm_cpu[TOPI][Relay] Add conv2d NHWC hybrid schedule for arm_cpuNov 16, 2023

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the changes @Anndrey24, overall LGTM! Just noticed something I missed previously

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated

@lhutton1lhutton1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

Implemented an `arm_cpu` conv2d NHWC schedule for fp32 using a hybrid GeMM approach, effectively breaking down the matrix multiplication into a macro-kernel (partitioning into fixed-sized, tile-level subproblems) and a micro-kernel (independently dealing with each subproblem). After the im2col transformation, the input matrix is handled natively (not interleaved), while the weights matrix is tiled and interleaved at compile time.
The micro-kernel uses 16 registers to accumulate the results of each 4x16 output tile, cycling through the operands needed to compute them (from the input and weight matrices) in the remaining registers.
There are now two ways to transform the weights matrix for conv2d, which are detailed in `convolution.cc`:
* for fp32: tile, interleave
* for int8: tile, interleave, transpose
To maintain naming consistency across both of these implementations (transposed vs not transposed), all mentions of `tile_rows_B` or `tile_cols_B` have been changed to `tile_N` and `tile_K` respectively to denote the tiling size along each axis of the flattened B matrix. As usual, `N = out_channels` and `K = kernel_width * kernel_height * in_channels`.
I have also added a new conv2d NHWC fp32 test for both the `conv2d_nhwc_spatial_pack` and `conv2d_NHWC_fp32_hybrid` schedules.

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Anndrey24 that is a lot of great work! 🚀 Some minor comments...

Comment threadpython/tvm/relay/op/strategy/arm_cpu.py Outdated
Comment threadpython/tvm/relay/op/strategy/arm_cpu.py
@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@ekaldaekalda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks great, thanks @Anndrey24!

@lhutton1

Copy link
Copy Markdown
Contributor

@tvm-bot rerun

@lhutton1
lhutton1 merged commit f38dc14 into apache:mainNov 24, 2023
@lhutton1

Copy link
Copy Markdown
Contributor

Thanks @Anndrey24@ekalda!

@Anndrey24
Anndrey24 deleted the fp32-hybrid-schedule branch November 24, 2023 15:56
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 17, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 22, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in apache#16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen test. The new `rewrite_simplify` rules are also covered by additional test cases.
lhutton1 pushed a commit that referenced this pull request Apr 24, 2024
This commit adds an `arm_cpu` conv2d NHWC schedule which generates SVE instructions by extending the hybrid GeMM approach implemented in #16106 to use scalable expressions as splitting factors.
Various vscale-related fixes needed to implement the schedule are also included, such as:
- adding vscale bounds in the `ConstIntBoundAnalyzer` and `IntervalSetEvaluator`
- simplifying `MinNode` and `MaxNode` that have scalable expression operands in `RewriteSimplifier`, which would appear when defining the shape of a buffer padded to be a multiple of vscale and in its respective buffer access indices (e.g. `C_1 = T.Buffer((1024 * (T.vscale() * 16 + 256 - 16 % T.vscale() * 16),), data=C)` instead of `C_1 = T.Buffer((1024 * (T.max(255, T.vscale() * 16 + 255 - 16 % T.vscale() * 16) + 1),), data=C)`)
The correctness of the new schedule is checked using a TOPI test, while the presence of generated SVE instructions is verified by a codegen_aarch64 test. The new rewrite_simplify rules are also covered by additional test cases.
Anndrey24 added a commit to Anndrey24/tvm that referenced this pull request Apr 29, 2024
…pu` targets
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in apache#16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in apache#16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
lhutton1 pushed a commit that referenced this pull request Apr 29, 2024
…pu` targets (#16951)
This patch partly reverts the unification of scalable and non-scalable scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
The non-scalable schedule for float32 splits the N axis (corresponding to number of output channels) by 16 in both the unified and the nonunified schedule versions, and then additionally splits the inner partitions by 4 in only the nonunified version to which this patch is reverting (first added in #16106). The two versions' behaviour would be equivalent if none of the padding on the N axis was removed during lowering, however we allow for that to happen as it proved to increase performance for very small convolutions.
As it stands, there seems to be a regression in cases where the datatype is float32 and the number of output channels is greater than 16, a multiple of 4, and not a multiple of 16, because even with the removed padding the nonunified schedule is able to vectorise over 4 elements, while the unified version cannot vectorise over 16 elements anymore.
Since all of the conv2d NHWC hybrid topi test cases used numbers of output channels either less than 16 or divisible by 16, this patch also adds a new case which falls in the aforementioned regression area.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Anndrey24@lhutton1@ekalda