Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

NVIDIA Deep Learning Examples for Tensor Cores

Introduction

This repository provides State-of-the-Art Deep Learning examples that are easy to train and deploy, achieving the best reproducible accuracy and performance with NVIDIA CUDA-X software stack running on NVIDIA Volta, Turing and Ampere GPUs.

NVIDIA GPU Cloud (NGC) Container Registry

These examples, along with our NVIDIA deep learning software stack, are provided in a monthly updated Docker container on the NGC container registry (https://ngc.nvidia.com). These containers include:

  • The latest NVIDIA examples from this repository
  • The latest NVIDIA contributions shared upstream to the respective framework
  • The latest NVIDIA Deep Learning software libraries, such as cuDNN, NCCL, cuBLAS, etc. which have all been through a rigorous monthly quality assurance process to ensure that they provide the best possible performance
  • Monthly release notes for each of the NVIDIA optimized containers

Computer Vision

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
ResNet-50PyTorchYesYesYes-Yes-YesYes-
ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
SE-ResNeXt-101PyTorchYesYesYes-Yes-YesYes-
EfficientNet-B0PyTorchYesYesYes----Yes-
EfficientNet-B4PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B0PyTorchYesYesYes----Yes-
EfficientNet-WideSE-B4PyTorchYesYesYes----Yes-
Mask R-CNNPyTorchYesYesYes-----Yes
nnUNetPyTorchYesYesYes----Yes-
SSDPyTorchYesYesYes-----Yes
ResNet-50TensorFlowYesYesYes----Yes-
ResNeXt101TensorFlowYesYesYes----Yes-
SE-ResNeXt-101TensorFlowYesYesYes----Yes-
Mask R-CNNTensorFlowYesYesYes----Yes-
SSDTensorFlowYesYesYes----YesYes
U-Net IndTensorFlowYesYesYes----YesYes
U-Net MedTensorFlowYesYesYes----Yes-
U-Net 3DTensorFlowYesYesYes----Yes-
V-Net MedTensorFlowYesYesYes----Yes-
U-Net MedTensorFlow2YesYesYes----Yes-
Mask R-CNNTensorFlow2YesYesYes----Yes-
EfficientNetTensorFlow2YesYesYesYes---Yes-
ResNet-50MXNet-YesYes------

Natural Language Processing

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
BERTPyTorchYesYesYesYes--YesYes-
TransformerXLPyTorchYesYesYesYes---Yes-
GNMTPyTorchYesYesYes------
TransformerPyTorchYesYesYes------
ELECTRATensorFlow2YesYesYesYes---Yes-
BERTTensorFlowYesYesYesYesYes-YesYesYes
BERTTensorFlow2YesYesYesYes---Yes-
BioBertTensorFlowYesYesYes----YesYes
TransformerXLTensorFlowYesYesYes------
GNMTTensorFlowYesYesYes------
Faster TransformerTensorflow----Yes----

Recommender Systems

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
DLRMPyTorchYesYesYes--YesYesYesYes
DLRMTensorFlow2YesYesYesYes---Yes-
NCFPyTorchYesYesYes------
Wide&DeepTensorFlowYesYesYes----Yes-
Wide&DeepTensorFlow2YesYesYes----Yes-
NCFTensorFlowYesYesYes----Yes-
VAE-CFTensorFlowYesYesYes------

Speech to Text

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
JasperPyTorchYesYesYes-YesYesYesYesYes
Hidden Markov ModelKaldi--Yes---Yes--

Text to Speech

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
FastPitchPyTorchYesYesYes----Yes-
FastSpeechPyTorch-YesYes-Yes----
Tacotron 2 and WaveGlowPyTorchYesYesYes-YesYesYesYes-

Graph Neural Networks

ModelsFrameworkA100AMPMulti-GPUMulti-NodeTRTONNXTritonDLCNB
SE(3)-TransformerPyTorchYesYesYes------

NVIDIA support

In each of the network READMEs, we indicate the level of support that will be provided. The range is from ongoing updates and improvements to a point-in-time release for thought leadership.

Glossary

Multinode Training
Supported on a pyxis/enroot Slurm cluster.

Deep Learning Compiler (DLC)
TensorFlow XLA and PyTorch JIT and/or TorchScript

Accelerated Linear Algebra (XLA)
XLA is a domain-specific compiler for linear algebra that can accelerate TensorFlow models with potentially no source code changes. The results are improvements in speed and memory usage.

PyTorch JIT and/or TorchScript
TorchScript is a way to create serializable and optimizable models from PyTorch code. TorchScript, an intermediate representation of a PyTorch model (subclass of nn.Module) that can then be run in a high-performance environment such as C++.

Automatic Mixed Precision (AMP)
Automatic Mixed Precision (AMP) enables mixed precision training on Volta, Turing, and NVIDIA Ampere GPU architectures automatically.

TensorFloat-32 (TF32)
TensorFloat-32 (TF32) is the new math mode in NVIDIA A100 GPUs for handling the matrix math also called tensor operations. TF32 running on Tensor Cores in A100 GPUs can provide up to 10x speedups compared to single-precision floating-point math (FP32) on Volta GPUs. TF32 is supported in the NVIDIA Ampere GPU architecture and is enabled by default.

Jupyter Notebooks (NB)
The Jupyter Notebook is an open-source web application that allows you to create and share documents that contain live code, equations, visualizations and narrative text.

Feedback / Contributions

We're posting these examples on GitHub to better support the community, facilitate feedback, as well as collect and implement contributions using GitHub Issues and pull requests. We welcome all contributions!

Known issues

In each of the network READMEs, we indicate any known issues and encourage the community to provide feedback.

About

Deep Learning Examples

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages