Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

346 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GPU4S Benchmark Suite documentation

Introduction

GPU4S is a benchmarking suite that rely on OBPMark-Kernel to test perfomance and reliability of GPUs and multi-threaded processors for space applications.

The benchmarking suites have been developed in order to be compatible with heterogeneous platforms:

  • Computer (x86_64, x86_32)
  • Android (ARM64, ARM32)
  • Nvidia Xavier/TX2 (ARM64) (currently in development)
List of tested devices Up to this date, the suite has been successfully compiled and executed on the following devices:
PlatformOperating SystemCPUGPUFrameworks Tested
High-end Laptop (x86_64)FedoraAMD Ryzen 7 7840HSNVIDIA GeForce RTX 4060 LaptopCPU, OpenMP, CUDA, OpenCL, HIP
Smartphone (ARM64)Android 5Qualcomm Snapdragon 810Adreno 430CPU, OpenMP, OpenCL

The benchmark uses a couple of different programming languages, libraries and frameworks to be able to compare perfomance of the same benchmark across most of devices:

  • Standard C/C++
  • CUDA
  • HIP
  • OpenCL
  • OpenMP

Background

Embedded GPUs have been identified by both private companies and government space agencies as a promising technology to meet the growing demands of payload processing. The GPU4S (GPU for Space) project, funded by the European Space Agency (ESA), explores the feasibility and benefits of using embedded GPUs for space workloads, and provides guidelines for their adoption in space applications.

Benchmark List and Basic Description

For most of the benchmark suite there is a naïve, optimized and library version. The benchmarks with their implementations are listed below.

BenchmarkNaïveOptimizedLibrary
Cifar 10✅ (CUDA only)
Cifar 10 Multiple✅ (CUDA only)
Convolution 2D✅ (CUDA only)
Correlation 2D
Fast Fourier Transform 2D
Fast Fourier Transform
Fast Fourier Transform Window
Finite Impulse Response Filter
Local Response Normalization (LRN)✅ (CUDA only)
Matrix Multiplication
Max Pooling✅ (CUDA only)
Memory Bandwidth
ReLU✅ (CUDA only)
Softmax✅ (CUDA only)
Wavelet Transform

Quick Start

If you already have the basic C/C++ programming tools installed (GCC/Clang, CMake ≥ 3.24, Git — see Prerequisites if not), you can try to compile and run the CPU version of the matrix multiplication benchmark in 3 steps:

# 1. Go to the benchmark directorycd gpu4s_benchmark/matrix_multiplication_bench
# 2. Generate build files and compile the matrix_mult CPU target
cmake -B build
cmake --build build --target cpu -j$(nproc)# 3. Run it
./build/bin/matrix_mult_cpu -s 512 -t
  • -s 512 runs the benchmark on a 512x512 matrix
  • -t prints the execution time

Congratulations! You have successfully built and run your first GPU4S benchmark.

Wanting to build with CUDA, HIP, OpenCL, or for Android? See docs/INSTALL.md for prerequisites and docs/BUILD_AND_RUN.md for building targets and check the runtime options.

Road map

Main focus:

  • fix issue of correctness between cpu and gpu in some benchmarks
  • refactor to extract the common of frameworks
  • fix of the clock to be executed in runtime + kernerCLK->deviceOBJ
  • Check for cl error during memory copy to host and clean
  • add UMA implementation for Android
  • Big cmake to compile everything
  • be compatible with jetson board + add UMA for jetson board

Bonus:

  • create a test with vulkan for android to have best performance (with softmax ?)
  • add map for verification of the result (ANDROID UMA)
  • refactor main, cuda, hip, opencl, OpenMP, -> create common
  • refactor cpu function -> create common (most important and easier)
  • add a get elapsed time function that print in this cpu function

The Authors

  • Ivan Rodriguez Ferrandez (BSC-UPC)
  • Alvaro Jover-Alvarez (BSC-UPC)
  • Leonidas Kosmidis (BSC-UPC)
  • Noah Perret (BSC-Centrale Nantes)
  • David Steenari (ESA)

License

ESA-PL Strong Copyleft – v2.5

About

Port GPU4S Benchmark suite to support cross-compilation for ARM embedded devices.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages