Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

rocHPL

rocHPL is a benchmark based on the HPL benchmark application, implemented on top of AMD's Radeon Open Compute ROCm Platform, runtime, and toolchains. rocHPL is created using the HIP programming language and optimized for AMD's latest discrete GPUs.

Requirements

Quickstart rocHPL build and install

Install script

You can build rocHPL using the install.sh script

# Clone rocHPL using git
git clone https://github.com/ROCm/rocHPL.git
# Go to rocHPL directory
cd rocHPL
# Run install.sh script
# Command line options:
# -h|--help - prints this help message
# -g|--debug - Set build type to Debug (otherwise build Release)
# --prefix=<dir> - Path to rocHPL install location (Default: build/rocHPL)
# --with-rocm=<dir> - Path to ROCm install (Default: /opt/rocm)
# --with-rocblas=<dir> - Path to rocBLAS library (Default: /opt/rocm/rocblas)
# --with-mpi=<dir> - Path to external MPI install (Default: clone+build OpenMPI)
# --with-mpi-gtl=<dir> - Path to external MPI-GTL install (Optional: defaults to no gtl support)
# --verbose-print - Verbose output during HPL setup (Default: true)
# --progress-report - Print progress report to terminal during HPL run (Default: true)
# --detailed-timing - Record detailed timers during HPL run (Default: true)
# --enable-tracing - Annotate profiler traces with rocTX markers (Default: false)
./install.sh

By default, UCX v1.18.0, and OpenMPI v5.0.7 will be cloned and built in rocHPL/tpl. After building, the rochpl executable is placed in build/rochpl-install.

Running rocHPL benchmark application

rocHPL provides some helpful wrapper scripts. A wrapper script for launching via mpirun is provided in mpirun_rochpl. This script has two distinct run modes:

mpirun_rochpl -P <P> -Q <P> -N <N> --NB <NB> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# N - is the total number of rows/columns of the global matrix
# NB - is the panel size in the blocking algorithm
# frac - is the split-update fraction (imporant for hiding some MPI
communication)

This run script will launch a total of np=PxQ MPI processes.

The second runmode takes an input file together with a number of MPI processes:

mpirun_rochpl -P <p> -Q <q> -i <input> -f <frac>
# where
# P - is the number of rows in the MPI grid
# Q - is the number of columns in the MPI grid
# input - is the input filename (default HPL.dat)
# frac - is the split-update fraction (important for hiding some MPI
communication)

The input file accpted by the rochpl executable follows the format below:

HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out output file name (if any)
0 device out (6=stdout,7=stderr,file)
1 # of problems sizes (N)
193280 Ns
1 # of NBs
896 NBs
1 PMAP process mapping (0=Row-,1=Column-major)
1 # of process grids (P x Q)
1 Ps
1 Qs
16.0 threshold
1 # of panel fact
2 PFACTs (0=left, 1=Crout, 2=Right)
1 # of recursive stopping criterium
32 NBMINs (>= 1)
1 # of panels in recursion
2 NDIVs
1 # of recursive panel fact.
2 RFACTs (0=left, 1=Crout, 2=Right)
1 # of broadcast
0 BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1 # of lookahead depth
1 DEPTHs (>=0)
1 SWAP (0=bin-exch,1=long,2=mix)
64 swapping threshold
0 L1 in (0=transposed,1=no-transposed) form
0 U in (0=transposed,1=no-transposed) form
0 Equilibration (0=no,1=yes)
8 memory alignment in double (> 0)

The mpirun_rochpl wraps a second script, run_rochpl, wherein some CPU core bindings are determined autmotically based on the node-local MPI grid.

Users wishing to launch rocHPL via a workload manager such as slurm may launch the run_rochpl script, or may launch the rochpl binary directly and specify CPU+GPU bindings via the job manager. For example:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -N 128000 -NB 512

When launching to multiple compute nodes, it can be useful to specify the local MPI grid layout on each node. To specify this, the -p and -q input parameters are used. For example, the srun line above is launching to two compute nodes, each with 8 GPUs. The local MPI grid layout can be specifed as either:

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 2 -q 4 -N 128000 -NB 512

or

srun -N 2 -n 16 -c 16 --gpus-per-task 1 --gpu-bind=closest ./build/bin/rochpl -P 4 -Q 4 -p 4 -q 2 -N 128000 -NB 512

This helps to control where/how much inter-node communication is occuring.

Performance evaluation

rocHPL is typically weak scaled so that the global matrix fills all available VRAM on all GPUs. The matrix size N is usually selected to be a multiple of the blocksize NB. Some sample runs on 288GB MI355X GPUs include:

  • 1 MI355X: mpirun_rochpl -P 1 -Q 1 -N 193280 --NB 896
  • 2 MI355X: mpirun_rochpl -P 2 -Q 1 -N 273280 --NB 896
  • 4 MI355X: mpirun_rochpl -P 2 -Q 2 -N 387072 --NB 896
  • 8 MI355X: mpirun_rochpl -P 2 -Q 4 -N 546560 --NB 896

Overall performance of the benchmark is measured in 64-bit floating point operations (FLOPs) per second. Performance is reported at the end of the run to the user's specified output (by default the performance is printed to stdout and a results file HPL.out).

See the Wiki for some common run configurations for various AMD Instinct GPUs.

Testing rocHPL

At the end of each benchmark run, residual error checking is computed, and PASS or FAIL is printed to output.

The simplest suite of tests should run configurations from 1 to 4 GPUs to exercise different communcation code paths. For example the tests:

mpirun_rochpl -P 1 -Q 1 -N 45312
mpirun_rochpl -P 1 -Q 2 -N 45312
mpirun_rochpl -P 2 -Q 1 -N 45312
mpirun_rochpl -P 2 -Q 2 -N 45312

should all report PASSED.

Please note that for successful testing, a device with at least 16GB of device memory is required.

Support

Please use the issue tracker for bugs and feature requests.

License

The license file can be found in the main repository.

About

High Performance Linpack for Next-Generation AMD HPC Accelerators

Resources

Stars

73 stars

Watchers

16 watching

Forks

Releases

Contributors

Languages