[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[PyTorch] Build custom ORT ops before running ONNX export tests - #1252

Merged
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops
Oct 16, 2024
Merged

[PyTorch] Build custom ORT ops before running ONNX export tests#1252
timmoon10 merged 4 commits into
NVIDIA:mainfrom
timmoon10:custom-ort-ops

Conversation

@timmoon10

@timmoon10timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
Member

Description

We have experienced failures in the ONNX export tests when running with Python 3.12 because PyPI does not have an available distribution for ONNX Runtime 1.13.1. Instead of manually rebuilding the custom ONNX Runtime ops at libcustom_ort_fp8_qdq_ops.so, I figure it's a good time to add logic to build the ops automatically before running the test.

Related: #41

Pinging @nzmora-nvidia and @asfiyab-nvidia.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refractor
  • Testing

Changes

  • Upgrade to ONNX Runtime 1.19.2 for ONNX export tests
  • Remove binary for custom ORT ops
  • Add code to build custom ORT ops
  • Export ONNX ops that perform intermediate compute in FP32, similar to TE kernels

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10timmoon10 added bug Something isn't working testing Improvements to tests or testing infrastructure labels Oct 15, 2024
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

Comment threadqa/L0_pytorch_unittest/test.sh

@cyanguwacyanguwa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
@timmoon10

Copy link
Copy Markdown
MemberAuthor

/te-ci pytorch

@timmoon10

timmoon10 commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm not sure what changed when I bumped the ONNX Runtime version, but I've experienced test failures for some FP8 operations (GeLU, LayerNorm, RMSNorm). Our tests tolerances basically require bit-perfect FP8 casting. However, this is not reasonable since the exported ONNX ops might do intermediate compute in FP16 while the TE kernels do all intermediate compute in FP32. I've modified the ONNX export infrastructure to match TE and do intermediate compute in FP32. In the future we should consider loosening the numerical tolerances to handle expected numerical error with FP8.

@ksivaman

Copy link
Copy Markdown
Member

@timmoon10 I haven't looked but there are some differences between the tests/pytorch/custom_ort_ops/custom_op_library.cc here and original version authored by @nzmora-nvidia from where we've copied this, could this be a source of some discrepancies?

@timmoon10

Copy link
Copy Markdown
MemberAuthor

Perhaps, but I don't think any of those changes should have affected numerics. In any case, the changes in this PR make the ONNX export more correct, so I don't think there's much risk to merging if the tests pass.

@timmoon10
timmoon10 merged commit f6b766b into NVIDIA:mainOct 16, 2024
timmoon10 added a commit to timmoon10/TransformerEngine that referenced this pull request Nov 7, 2024
…IA#1252)
* Build custom ORT ops before running ONNX tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Remove ONNX from context parallelism tests
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Export ONNX ops that do compute in FP32
Matches internal impl of TE kernels.
Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Add build script for custom ORT ops
Signed-off-by: Tim Moon <tmoon@nvidia.com>
---------
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't workingtestingImprovements to tests or testing infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@timmoon10@ksivaman@cyanguwa