Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [PyTorch] Use Quantization API for reference NVFP4 recipe by negvet · Pull Request #2259 · NVIDIA/TransformerEngine · GitHub
Skip to content

[PyTorch] Use Quantization API for reference NVFP4 recipe - #2259

Merged
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api
Oct 14, 2025
Merged

[PyTorch] Use Quantization API for reference NVFP4 recipe#2259
negvet merged 6 commits into
NVIDIA:mainfrom
negvet:merge_middleware_quantization_api

Conversation

@negvet

@negvetnegvet commented Oct 10, 2025

Copy link
Copy Markdown
Collaborator

Description

This PR uses quantization abstractions from #2039.

This is just a first refactor step.
TODO in the upcoming PRs: decouple python API, create ref recipe zoo, ops define roles, use for attention, etc.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

negvetand others added 5 commits October 2, 2025 12:31
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
Signed-off-by: Evgeny <etsykunov@nvidia.com>
@negvet

Copy link
Copy Markdown
CollaboratorAuthor

/te-ci pytorch

@ptrendx
ptrendx requested a review from CopilotOctober 10, 2025 16:07

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refactors the experimental NVFP4 quantization code to use the new quantization abstractions from TransformerEngine. The refactoring migrates from custom experimental classes to the standard quantization API while maintaining the reference implementation functionality.

  • Removes the get_module_quantizers function and uses direct quantizer methods in modules
  • Consolidates experimental classes into the standard quantization API
  • Replaces environment-based configuration with factory-based quantizer creation

Reviewed Changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated no comments.

Show a summary per file
FileDescription
transformer_engine/pytorch/module/linear.pyUpdates to use direct quantizer methods instead of get_module_quantizers
transformer_engine/pytorch/module/layernorm_linear.pySame quantizer method updates as linear module
transformer_engine/pytorch/module/_common.pyRemoves deprecated get_module_quantizers function and related imports
transformer_engine/pytorch/experimental/quantization_nvfp4.pyMajor refactor moving from experimental to standard API with new factory function
transformer_engine/pytorch/experimental/quantization.pyRemoves experimental base classes that were moved to standard API
transformer_engine/pytorch/experimental/gemm.pyUpdates to use standard quantization tensor storage
transformer_engine/pytorch/experimental/config.pyCompletely removed - functionality replaced by factory pattern
transformer_engine/pytorch/experimental/init.pyRemoves exports from deleted config module
transformer_engine/pytorch/attention/dot_product_attention/dot_product_attention.pyAdds early return for custom recipes
tests/pytorch/nvfp4/*.pyUpdates test imports and replaces environment-based config with factory-based approach
tests/pytorch/distributed/run_numerics_exact.pySame test configuration updates as other test files
Comments suppressed due to low confidence (1)

transformer_engine/pytorch/experimental/quantization_nvfp4.py:1

  • Removed call to dequantize method in repr but the method was also removed from the class, which could cause AttributeError if this line was still present.
# Copyright (c) 2022-2025, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@timmoon10timmoon10 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@negvet
negvet merged commit dfacd9f into NVIDIA:mainOct 14, 2025
22 of 23 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@negvet@timmoon10