Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Getting started

The purpose of these examples is to help demonstrates various transfer and timing techniques.

Transfer techniques

  1. Manual cudaMemcpy
  2. Manual cudaMemcpyAsync
  3. Unified memory
  4. Unified memory with cudaMemPrefetchAsync

Timing techniques

  1. Chrono
  2. CUDA events
  3. NVTX markers

Usage

Each test does the follow:

  1. Transfers two chunks of data to the GPU.
  2. Run simple kernel on each chunk (mutually exclusive)
  3. Transfer both chunks back to CPU.
  4. Verify results

Throughput based on entire workflow.

Default transfer size is 1GB.

./cudaMemcpyAsync
Running with = 1073741824 B (1.07 GB)
Chrono: 194.463501 ms @ 5.521560 GB/s
Events: 195.145355 ms @ 5.502267 GB/s

or add number of values to transfer

./cudaMemcpyAsync 1000000000
Running with = 4000000000 B (4.00 GB)
Chrono: 725.302307 ms @ 5.514942 GB/s
Events: 726.420593 ms @ 5.506452 GB/s

Using NVTX markers

We must use Nsight Systems to see NVTX. Open *.qdrep file with Nsight Systems GUI.

nsys profile -s none -t cuda,nvtx --stats=true ./cudaMemcpyAsync
WARNING: Backtraces will not be collected because sampling is disabled.
Collecting data...
Running with = 1073741824 B (1.07 GB)
Chrono: 195.130112 ms @ 5.502697 GB/s
Events: 195.151413 ms @ 5.502096 GB/s
Processing events...
Capturing symbol files...
Saving temporary "/tmp/nsys-report-8b68-3d6e-a890-0843.qdstrm" file to disk...
Creating final output files...
Processing [==============================================================100%]
Saved report file to "/tmp/nsys-report-8b68-3d6e-a890-0843.qdrep"
Exporting 1532 events: [==================================================100%]
Exported successfully to
/tmp/nsys-report-8b68-3d6e-a890-0843.sqlite
CUDA API Statistics:
Time(%) Total Time (ns) Num Calls Average Minimum Maximum Name ------- --------------- --------- ------------- ----------- ----------- ------------------------
92.8 7,935,281,846 30 264,509,394.9 162,979,208 367,487,697 cudaStreamSynchronize 3.8 328,299,372 2 164,149,686.0 164,072,205 164,227,167 cudaHostAlloc 1.9 158,486,349 2 79,243,174.5 79,163,730 79,322,619 cudaFreeHost 1.5 130,534,008 2 65,267,004.0 701,703 129,832,305 cudaMalloc 0.0 1,454,937 2 727,468.5 691,604 763,333 cudaFree 0.0 401,077 60 6,684.6 2,374 26,921 cudaMemcpyAsync 0.0 373,116 30 12,437.2 4,637 31,135 cudaLaunchKernel 0.0 83,993 10 8,399.3 5,491 10,170 cudaEventRecord 0.0 18,491 5 3,698.2 3,500 3,931 cudaEventSynchronize 0.0 7,857 2 3,928.5 873 6,984 cudaStreamCreate 0.0 7,470 2 3,735.0 1,390 6,080 cudaStreamDestroy 0.0 5,474 2 2,737.0 397 5,077 cudaEventCreateWithFlags
0.0 2,602 2 1,301.0 417 2,185 cudaEventDestroy CUDA Kernel Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Name ------- --------------- --------- ----------- --------- --------- -----------------------------------------------------------------------
50.4 67,362,653 15 4,490,843.5 4,469,005 4,518,252 void VectorOperation<Add<float>, float>(int, float, float*, Add<float>)
49.6 66,399,104 15 4,426,606.9 4,391,821 4,462,860 void VectorOperation<Sub<float>, float>(int, float, float*, Sub<float>)
CUDA Memory Operation Statistics (by time):
Time(%) Total Time (ns) Operations Average Minimum Maximum Operation ------- --------------- ---------- ------------- ----------- ----------- ------------------
50.8 5,314,309,114 30 177,143,637.1 162,980,699 191,320,991 [CUDA memcpy DtoH]
49.2 5,142,402,501 30 171,413,416.7 168,996,768 173,680,811 [CUDA memcpy HtoD]
CUDA Memory Operation Statistics (by size in KiB):
Total Operations Average Minimum Maximum Operation -------------- ---------- ------------- ------------- ------------- ------------------
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy DtoH]
31,457,280.000 30 1,048,576.000 1,048,576.000 1,048,576.000 [CUDA memcpy HtoD]
NVTX Push-Pop Range Statistics:
Time(%) Total Time (ns) Instances Average Minimum Maximum Range ------- --------------- --------- ------------- ----------- ----------- ------------
67.6 3,889,587,729 5 777,917,545.8 777,031,465 778,565,183 Process_Loop
21.7 1,247,130,464 5 249,426,092.8 247,611,878 250,019,880 Verify 10.7 616,467,941 5 123,293,588.2 120,541,180 132,044,582 Reset 0.0 101,082 5 20,216.4 16,613 22,645 H2D_A 0.0 93,874 5 18,774.8 18,020 21,060 Kernel_A 0.0 29,242 5 5,848.4 5,317 6,468 Kernel_B 0.0 23,568 5 4,713.6 4,374 4,970 D2H_A 0.0 18,801 5 3,760.2 3,283 4,147 H2D_B 0.0 14,849 5 2,969.8 2,680 3,182 D2H_B Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.qdrep"
Report file moved to "/home/belt/workStuff/git_examples/transfer_examples/report4.sqlite"

About

Various CPU<->GPU transfer and timing techniques

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages