Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); [PyTorch] Prune L0 unit test by ksivaman · Pull Request #1999 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Prune L0 unit test - #1999

Merged
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests
Jul 29, 2025
Merged

[PyTorch] Prune L0 unit test#1999
ksivaman merged 8 commits into
NVIDIA:mainfrom
ksivaman:prune_pytorch_L0_tests

Conversation

@ksivaman

Copy link
Copy Markdown
Member

Description

Reduce L0 unit testing runtime without removing significant coverage.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

  • Print detailed logs only for failed tests to reduce logging overheads.
  • Initialize supported recipes beforehand instead of launching multiple tests and then skipping unsupported options.
  • Remove unused imports.
  • Remove redundant/duplicate tests. E.g. remove cpu offloading and CG tests from sanity as they are individually tested more exhaustively in their own L0 scripts.
  • For slow tests such as ONNX export, refactor the tests to split the @parametrized options into individual tests instead of combining them all together, which leads to exponentiation. This way, one feature is tested at a time while setting reasonable defaults for others. For the recipe specifically, the order of recipes in the list ensures different defaults on different devices based on available support to maximize coverage.

Maybe in the future

  • Distributed/other tests.
  • Apply the split test refactor to other tests if they grow; currently most tests run in under 3 mins without the logging overhead.
  • Test one recipe only per device?
  • In case of rapid growth, have 2 levels of sanity, L0 and L1.

Local testing times (mins)

  • H100: 34 → 14
  • B100: 39 → 13

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivamanksivaman added the testing Improvements to tests or testing infrastructure label Jul 27, 2025
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@ksivaman
ksivaman merged commit aac7442 into NVIDIA:mainJul 29, 2025
21 checks passed
ksivaman added a commit to ksivaman/TransformerEngine-1 that referenced this pull request Jul 29, 2025
* Add verbosity only for failing tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune some tests and preinit recipe
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Prune further tests
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* fix multitensor
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* Minor fixes
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix a100
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
---------
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ksivaman@cyanguwa