Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - dbeer/nvbandwidth: A tool for bandwidth measurements on NVIDIA GPUs. · GitHub
Skip to content

Repository files navigation

nvbandwidth

A tool for bandwidth measurements on NVIDIA GPUs.

Measures bandwidth for various memcpy patterns across different links using copy engine or kernel copy methods. nvbandwidth reports current measured bandwidth on your system. Additional system-specific tuning may be required to achieve maximal peak bandwidth.

Requirements

nvbandwidth requires the installation of a CUDA toolkit and some additional Linux software components to be built and run. This section provides the relevant details Install a cuda toolkit (version 11.X or above)

Install a compiler package which supports c++17. GCC 7.x or above is a possible option.

Install cmake (version 3.20 or above)

Install Boost program options library (More details in the next section)

Ensure that path to nvcc binary (install via toolkit) is available in the $PATH variable on linux systems

Dependencies

To build and run nvbandwidth please install the Boost program_options library (https://www.boost.org/doc/libs/1_66_0/doc/html/program_options.html).

Ubuntu/Debian users can run the following to install:

apt install libboost-program-options-dev

On Ubuntu/Debian, we have provided a utility script (debian_install.sh) which installs some generic software components needed for the build. The script also builds the nvbandwidth project.

sudo ./debian_install.sh

Fedora users can run the following to install:

sudo dnf -y install boost-devel

Build

To build the nvbandwidth executable:

cmake .
make

You may need to set the BOOST_ROOT environment variable on Windows to tell CMake where to find your Boost installation.

Usage:

./nvbandwidth -h
nvbandwidth CLI:
-h [ --help ] Produce help message
-b [ --bufferSize ] arg (=64) Memcpy buffer size in MiB
-l [ --list ] List available testcases
-t [ --testcase ] arg Testcase(s) to run (by name or index)
-v [ --verbose ] Verbose output
-s [ --skipVerification ] Skips data verification after copy
-d [ --disableAffinity ] Disable automatic CPU affinity control
-i [ --testSamples ] arg (=3) Iterations of the benchmark
-m [ --useMean ] Use mean instead of median for results

To run all testcases:

./nvbandwidth

To run a specific testcase:

./nvbandwidth -t device_to_device_memcpy_read_ce

Example output:

Running device_to_device_memcpy_write_ce.
memcpy CE GPU(row) <- GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 0.00 276.07 276.36 276.14 276.29 276.48 276.55 276.33
1 276.19 0.00 276.29 276.29 276.57 276.48 276.38 276.24
2 276.33 276.29 0.00 276.38 276.50 276.50 276.29 276.31
3 276.19 276.62 276.24 0.00 276.29 276.60 276.29 276.55
4 276.03 276.55 276.45 276.76 0.00 276.45 276.36 276.62
5 276.17 276.57 276.19 276.50 276.31 0.00 276.31 276.15
6 274.89 276.41 276.38 276.67 276.41 276.26 0.00 276.33
7 276.12 276.45 276.12 276.36 276.00 276.57 276.45 0.00

Set number of iterations and the buffer size for copies with --testSamples and --bufferSize

Test Details

There are two types of copies implemented, Copy Engine (CE) or Steaming Multiprocessor (SM)

CE copies use memcpy APIs. SM copies use kernels.

SM copies will truncate the copy size to fit uniformly on the target device to correctly report the bandwidth. The actual byte size for the copy is:

(threadsPerBlock * deviceSMCount) * floor(copySize / (threadsPerBlock * deviceSMCount))

threadsPerBlock is set to 512.

Measurement Details

A blocking kernel and CUDA events are used to measure time to perform copies via SM or CE, and bandwidth is calculated from a series of copies.

First, we enqueue a spin kernel that spins on a flag in host memory. The spin kernel spins on the device until all events for measurement have been fully enqueued into the measurement streams. This ensures that the overhead of enqueuing operations is excluded from the measurement of actual transfer over the interconnect. Next, we enqueue a start event, certain count of memcpy iterations, and finally a stop event. Finally, we release the flag to start the measurement.

This process is repeated 3 times, and the median bandwidth for each trial is reported.

Number of repetitions can be overriden using the --testSamples option, and in order to use arithmetic mean instead of median you can specify --useMean option.

Unidirectional Bandwidth Tests

Running host_to_device_memcpy_ce.
memcpy CE CPU(row) -> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 26.03 25.94 25.97 26.00 26.19 25.95 26.00 25.97

Unidirectional tests measure the bandwidth between each pair in the output matrix individually. Traffic is not sent simultaneously.

Bidirectional Host <-> Device Bandwidth Tests

Running host_to_device_bidirectional_memcpy_ce.
memcpy CE CPU(row) <-> GPU(column) bandwidth (GB/s)
0 1 2 3 4 5 6 7
0 18.56 18.37 19.37 19.59 18.71 18.79 18.46 18.61

The setup for bidirectional host to device bandwidth transfer is shown below:

Stream 0 (measured stream) performs writes to the device, while the interfering stream in the opposite direction produces reads. This pattern is reversed for measuring bidirectional device to host bandwidth as shown below.

Bidirectional Device <-> Device Bandwidth Tests

The setup for bidirectional device to device transfers is shown below:

The test launches traffic on two streams: stream 0 launched on device 0 performs writes from device 0 to device 1 while the interference stream 1 launches opposite traffic performing writes from device 1 to device 0.

CE bidirectional bandwidth tests calculate bandwidth on the measured stream:

CE bidir. bandwidth = (size of data on measured stream) / (time on measured stream)

However, SM bidirectional test launches memcpy kernels on source and peer GPUs as independent streams and calculates bandwidth as:

SM bidir. bandwidth = size/(time on stream1) + size/(time on stream2)

About

A tool for bandwidth measurements on NVIDIA GPUs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages