Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [PyTorch] Allocate grouped linear wgrads as tensor views by timmoon10 · Pull Request #3049 · NVIDIA/TransformerEngine · GitHub
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions transformer_engine/pytorch/csrc/extensions/allocate.cpp
Original file line numberDiff line numberDiff line change
Expand Up@@ -12,6 +12,16 @@
namespace transformer_engine {
namespace pytorch {

/* Allocate multiple PyTorch tensors backed by the same buffer.
*
* Use with caution and avoid exposing externally.
*
* In order to reduce CPU overhead, we compute pointer offsets
* manually and construct PyTorch tensors with raw pointers. The
* backing buffer is deallocated once the final tensor is destroyed.
* Stream usage is not recorded, so there may be race conditions if
* compute is performed on multiple streams.
*/
std::vector<at::Tensor> bulk_allocate(const std::vector<std::vector<size_t>> &shapes,
const std::vector<at::ScalarType> &dtypes,
std::optional<c10::Device> device,
Expand Down
12 changes: 6 additions & 6 deletions transformer_engine/pytorch/module/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -496,13 +496,13 @@ def backward(
if ctx.fuse_wgrad_accumulation:
wgrad_list = main_grads
else:
weight_shape = list(weights[0].size())
wgrad_list = tex.bulk_allocate(
[weight_shape] * ctx.num_gemms,
[ctx.activation_dtype] * ctx.num_gemms,
ctx.device,
[256] * ctx.num_gemms, # alignment
wgrad_packed = torch.empty(
ctx.num_gemms,
*weights[0].size(),
dtype=ctx.activation_dtype,
device=ctx.device,
)
wgrad_list = [wgrad_packed[i] for i in range(ctx.num_gemms)]
Comment thread
timmoon10 marked this conversation as resolved.

if ctx.save_original_input:
inp = inputmats[0]
Expand Down
10 changes: 5 additions & 5 deletions transformer_engine/pytorch/ops/basic/grouped_linear.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -1393,12 +1393,12 @@ def _fuser_backward_split_quantize(
]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
grad_weights = tex.bulk_allocate(
[weight_shape] * num_groups,
[ctx.dtype] * num_groups,
device,
[256] * num_groups, # alignment
grad_weights_packed = torch.empty(
grouped_shape,
dtype=ctx.dtype,
device=device,
)
grad_weights = [grad_weights_packed[i] for i in range(num_groups)]
Comment on lines +1396 to +1401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2Per-tensor 256-byte alignment no longer guaranteed

tex.bulk_allocate(..., [256] * num_groups) guaranteed each wgrad tensor started on a 256-byte boundary. With torch.empty(grouped_shape, ...), PyTorch aligns the base allocation but sub-tensors grad_weights_packed[i] land at byte offset i * out_features * in_features * element_size. That offset is a multiple of 256 bytes only when out_features * in_features * element_size % 256 == 0. For typical large weight matrices this holds, but it is not enforced. If a future model uses an unusual hidden dimension the misalignment could cause a silent performance regression in the GEMM kernel. The same pattern appears in backward_grouped_mlp.py and module/grouped_linear.py.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is true, there should be some check added for the sizes to be right.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is possible, but would balloon the surface area of this PR for little benefit. The proper fix:

  • Add a memoized dtype_bytes/dtype_bits utility function, probably in utils.py.
  • To handle future dtypes, dtype_bytes should support both torch.dtype and tex.DType.
  • utils.py now depends on transformer_engine_torch and is no longer purely upstream to the rest of the package.

It is somewhat sloppy and dangerous to just take views blindly, but even 16x8 BF16 weights will fulfill the alignment requirement.

final_weight_grads = list(grad_weights)

# Perform dgrad GEMMs
Expand Down
11 changes: 6 additions & 5 deletions transformer_engine/pytorch/ops/fused/backward_grouped_mlp.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -197,12 +197,13 @@ def _compute_grad_params(
w_list = [get_main_grad_from_param(w, op_label=op_label) for w in weights]
accumulate_into_main_grad = get_accumulate_flag_in_param(weights[0])
else:
w_list = tex.bulk_allocate(
[weight_shape] * num_groups,
[dtype] * num_groups,
device,
[256] * num_groups, # alignment
wgrad_packed = torch.empty(
num_groups,
*weight_shape,
dtype=dtype,
device=device,
)
w_list = [wgrad_packed[i] for i in range(num_groups)]
wgrad_output = w_list

if ctx.weight_requires_grad:
Expand Down
Loading