Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add ONNX export support for TE modules by asfiyab-nvidia · Pull Request #41 · NVIDIA/TransformerEngine · GitHub
Skip to content

Add ONNX export support for TE modules - #41

Merged
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support
Jan 18, 2023
Merged

Add ONNX export support for TE modules#41
ksivaman merged 17 commits into
NVIDIA:mainfrom
asfiyab-nvidia:dev-onnx-export-support

Conversation

@asfiyab-nvidia

Copy link
Copy Markdown
Contributor
  • Add TorchScript Operators
  • Add symbolic methods to ONNX exporter
  • Add tests for the ONNX export

Signed-off-by: Asfiya Baig asfiyab@nvidia.com
Signed-off-by: Neta Zmora nzmora@nvidia.com

@ptrendx

Copy link
Copy Markdown
Member

Hi @asfiyab-nvidia, what is that libcustom so file?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx it contains the onnxruntime (ORT) implementations for FP8 functionality. This is used to test the ONNX export and validate the ORT outputs against TE outputs. (code under tests/test_onnx_export.py)
The so is included in the PR so there's no dependencies on external sources

@ptrendx

Copy link
Copy Markdown
Member

Does that have to be closed source? If so, can we at least move it to tests directory instead of the top level one? If it does not have to be closed source then maybe we can have the source inside tests and compile it on the fly?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

Moving the .so to the tests directory seems to be a better approach at the moment. We can potentially include the source code in a follow up PR.

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

Comment threadtransformer_engine/pytorch/csrc/ts_fp8_op.cpp Outdated
Comment threadtransformer_engine/pytorch/module.py Outdated
@ptrendx

Copy link
Copy Markdown
Member

Please fix the tests (see the results for commit 4812408) - the biggest problem is that you try to run tests requiring FP8 on non-Hopper, which triggers the assertion failure. I am working on enabling Hopper GPU in CI, so we should be able to get the FP8 tests running soon too.

@netaz

netaz commented Jan 8, 2023

Copy link
Copy Markdown

@ptrendx is there some code in TE we can leverage to query the SM version, or do you recommend us installing some lib (e.g. pynvml)?

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

1 similar comment
@ptrendx

Copy link
Copy Markdown
Member

/te-ci

@asfiyab-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx can you please authorize a pipeline run for the latest commit? It contains fixes for the failures from the last run. Thanks

@ptrendx

Copy link
Copy Markdown
Member

/te-ci

asfiyab-nvidiaand others added 14 commits January 17, 2023 20:18
* Add TorchScript Operators
* Add symbolic methods to ONNX exporter
* Add tests for the ONNX export
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
* Increase layernorm FP16 threshold
* Normalize onnx file names: _ separates configs; - separates words in a single config
* Add get_attn_mask_str and fix mask string
* Add missing ONNX files
* Moved generated ONNX files to tests/gen_onnx_models/
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. remove List import for pylint failure
2. address comments: remove state tensors from GPU
3. address comments: Update reverse_map_dtype function and add to namespace
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. skip FP8 tests on non-hopper devices
2. minor fix for C++ lint check
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>
1. update copyrights
2. update path to ORT .so
Signed-off-by: Asfiya Baig <asfiyab@nvidia.com>

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initial comments

Comment threadtests/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/__init__.py Outdated
Comment threadtests/test_onnx_export.py Outdated
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: asfiyab-nvidia <117682710+asfiyab-nvidia@users.noreply.github.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

1 similar comment
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivamanksivaman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ksivaman
ksivaman merged commit 6c9ce17 into NVIDIA:mainJan 18, 2023
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
[New API] Added support for Reshape operation.
[New API] Added support for DgradDreluBNBwdWeight operation
[Minor Enhancement] Added cudnn frontend enums to simplify Resample operation creation.
[Minor Enhancement] Added alpha and beta values as key for the plan caches.
[Bug Fix] Fixed an error which was causing reference code to fail with segmentation fault.
[Bug Fix] Fixed an issue where stride/padding and dilation values were incorrectly cached for 2d convolutions.
[Bug Fix] Fixed issues where error statuses were not handled correctly during tensor creation.
[Samples] Added a new sample to show case how fMHA graph can be programmed through FE API. This sample contains both fprop and backprop graphs.
[Samples] Added a new sample to show case DgradDreluBNBwdWeight operation.
[Samples] Added a modular block which models fprop of residual block resnet.
Co-authored-by: Anerudhan Gopal <agopal@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@asfiyab-nvidia@ptrendx@netaz@ksivaman