Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Fix FlashAttention tests by tcherckez-nvidia · Pull Request #99 · NVIDIA/TransformerEngine · GitHub
Skip to content

Fix FlashAttention tests - #99

Merged
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests
Mar 29, 2023
Merged

Fix FlashAttention tests#99
ksivaman merged 2 commits into
NVIDIA:mainfrom
tcherckez-nvidia:tcherckez-fix-flash-attn-tests

Conversation

@tcherckez-nvidia

Copy link
Copy Markdown
Contributor

No description provided.

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 6606f27 to 7a441a8CompareMarch 14, 2023 11:50
Comment threadtests/pytorch/test_onnx_export.py Outdated
@ksivaman

Copy link
Copy Markdown
Member

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@tcherckez-nvidia

tcherckez-nvidia commented Mar 16, 2023

Copy link
Copy Markdown
ContributorAuthor

We current disable FA explicitly when running the onnx export tests in qa/L0_unittest/test.sh. This should also be removed in this PR.

@ksivaman
Should I then remove NVTE_FLASH_ATTN? I don't its use other then in disabling the usage of FA in tests

@ksivaman

Copy link
Copy Markdown
Member

@tcherckez-nvidia Yes

@ptrendx

Copy link
Copy Markdown
Member

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

@ptrendx

ptrendx commented Mar 17, 2023

Copy link
Copy Markdown
Member

Also, whichever path we choose, we should create a proper end-to-end workflow tutorial - from training to inference, including in-framework inference as well as export to TRT and FasterTransformer.

(Obviously, that is out of scope for this PR)

@nzmora-nvidia

nzmora-nvidia commented Mar 18, 2023

Copy link
Copy Markdown
Contributor

Maybe we can just have a global function, something like transformer_engine.pytorch.prepare_onnx_export() or something like that, which would then be checked by each module (which would then use the most basic implementation if it is set)?

Good idea @ptrendx. Perhaps in the form of a context manager like te.fp8_autocast()

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch 2 times, most recently from 07224ac to fba0d2fCompareMarch 19, 2023 11:53
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from fba0d2f to 343bcd1CompareMarch 20, 2023 13:23
Comment threadtests/pytorch/test_onnx_export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtransformer_engine/pytorch/export.py Outdated
Comment threadtransformer_engine/pytorch/transformer.py Outdated
Comment threadtests/pytorch/test_onnx_export.py Outdated
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 343bcd1 to 9162f00CompareMarch 21, 2023 08:03
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

/te-ci

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 9162f00 to 80211e2CompareMarch 21, 2023 08:04
@tcherckez-nvidia

Copy link
Copy Markdown
ContributorAuthor

@ptrendx@ksivaman please review

@ptrendx

Copy link
Copy Markdown
Member

@tcherckez-nvidia please resolve the merge conflicts.

@ptrendxptrendx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good.

@ptrendx

Copy link
Copy Markdown
Member

Also, could you add the new context manager to the pytorch API docs here https://github.com/NVIDIA/TransformerEngine/blob/main/docs/api/pytorch.rst?

@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 80211e2 to 7582c88CompareMarch 26, 2023 05:32
Comment threaddocs/api/pytorch.rst Outdated
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
@tcherckez-nvidia
tcherckez-nvidiaforce-pushed the tcherckez-fix-flash-attn-tests branch from 7582c88 to 1eccf06CompareMarch 26, 2023 07:53
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
@ksivaman

Copy link
Copy Markdown
Member

/te-ci

@ksivaman
ksivaman merged commit bcbd4be into NVIDIA:mainMar 29, 2023
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Mar 31, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
cyanguwa pushed a commit to cyanguwa/TransformerEngine that referenced this pull request Apr 1, 2023
Signed-off-by: Tal Cherckez <tcherckez@nvidia.com>
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: Charlene Yang <charleney@nvidia.com>
zhiyu-deep pushed a commit to zhiyu-deep/TransformerEngine that referenced this pull request Sep 3, 2024
- Bug fix
- Fixed an issue where custom dropout mask was not correctly applied.
- Added `-fvisibility=hidden` for the pip wheels generated to avoid
symbol conflicts with other modules that use cudnn frontend.
- Fixed an issue in sdpa kernels which will lead to numerical
mismatches.
- Fixed an issue in sdpa fp8 fprop kernels (in inference mode)
- Samples
- Added a new sample to showcase how a custom dropout mask can be
applied to a sdpa operation.
- Added a sample to shocase convolutions on large (`c * d * h * w > 2 **
31`) tensors.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tcherckez-nvidia@ksivaman@ptrendx@nzmora-nvidia