Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - fml-fam/fml: Fused Matrix Library · GitHub
Skip to content

Repository files navigation

fml

fml is the Fused Matrix Library, a multi-source, header-only C++ library for dense matrix computing. The emphasis is on real-valued matrix types (float, double, and __half) for numerical operations useful for data analysis.

The goal of fml is to be "medium-level". That is, high-level compared to working directly with e.g. the BLAS or CUDA™, but low(er)-level compared to other C++ matrix frameworks. Some knowledge of the use of LAPACK will make many choices in fml make more sense.

The library provides 4 main classes: cpumat, gpumat, parmat, and mpimat. These are mostly what they sound like, but the particular details are:

  • CPU: Single node cpu computing (multi-threaded if using multi-threaded BLAS and linking with OpenMP).
  • GPU: Single gpu computing.
  • MPI: Multi-node computing via ScaLAPACK (+gpus if using SLATE).
  • PAR: Multi-node and/or multi-gpu computing.

There are some differences in how objects of any particular type are constructed. But the high level APIs are largely the same between the objects. The goal is to be able to quickly create laptop-scale prototypes that are then easily converted into large scale gpu/multi-node/multi-gpu/multi-node+multi-gpu codes.

Installation

The library is header-only so no installation is strictly necessary. You can just include a copy/submodule in your project. However, if you want some analogue of make install, then you could do something like:

ln -s ./src/fml /usr/include/

Dependencies and Other Software

There are no external header dependencies, but there are some shared libraries you need to have (more information below):

Other software we use:

  • Tests use catch2 (a copy of which is included under tests/).

You can find some examples of how to use the library in the examples/ tree. Right now there is no real build system beyond some ad hoc makefiles; but ad hoc is better than no hoc.

Depending on which class(es) you want to use, here are some general guidelines for using the library in your own project:

  • CPU: cpumat
    • Compile with your favorite C++ compiler.
    • Link with LAPACK and BLAS (and ideally with OpenMP).
  • GPU: gpumat
    • Compile with nvcc.
    • For most functionality, link with libcudart, libcublas, and libcusolver. Link with libcurand if using the random generators. Link with libnvidia-ml if using nvml (if you're only using this, then you don't need nvcc; an ordinary C++ compiler will do). If you have CUDA installed and do not know what to link with, there is no harm in linking with all of these.
  • MPI: mpimat
    • Compile with mpicxx.
    • Link with libscalapack.
  • PAR: parmat
    • Compile with mpicxx.
    • Link with CPU stuff if using parmat_cpu; link with GPU stuff if using parmat_gpu (you can use both).

Check the makefiles in the examples/ tree if none of that makes sense.

Example

Here's a simple example computing the SVD with some data held on a single CPU:

#include<fml/cpu.hh>usingnamespacefml;intmain()
{
len_t m = 3;
len_t n = 2;
cpumat<float> x(m, n);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
s.info();
s.print();
return0;
}

Save as svd.cpp and build with:

g++ -I/path/to/fml/src -fopenmp svd.cpp -o svd -llapack -lblas

You should see output like

# cpumat 3x2 type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

The API is largely the same if we change the object storage, but we have to change the object initialization. For example, if x is an object of class mpimat, we still call linalg::svd(x, s). The differences lie in the creation of the objects. Here is how we might change the above example to use distributed data:

#include<fml/mpi.hh>usingnamespacefml;intmain()
{
grid g = grid(PROC_GRID_SQUARE);
g.info();
len_t m = 3;
len_t n = 2;
mpimat<float> x(g, m, n, 1, 1);
x.fill_linspace(1.f, (float)m*n);
x.info();
x.print(0);
cpuvec<float> s;
linalg::svd(x, s);
if (g.rank0())
{
s.info();
s.print();
}
g.exit();
g.finalize();
return0;
}

In practice, using such small block sizes for an MPI matrix is probably not a good idea; we only do so for the sake of demonstration (we want each process to own some data). We can build this new example via:

mpicxx -I/path/to/fml/src svd.cpp -fopenmp svd.cpp -o svd -lscalapack-openmpi

We can launch the example with multiple processes via

mpirun -np 4 ./svd

And here we see:

## Grid 0 2x2
# mpimat 3x2 on 2x2 grid type=f
1 4 2 5 3 6 # cpuvec 2 type=f
9.5080 0.7729 

High-Level Language Bindings

Header and API Stability

tldr:

  • Use the super headers (or read the long explanation)
    • CPU - fml/cpu.hh
    • GPU - fml/gpu.hh
    • MPI - fml/mpi.hh
    • PAR still evolving
  • Existing API's are largely stable. Most changes will be additions rather than modifications.

The project is young and things are still mostly evolving. The current status is:

Headers

There are currently "super headers" for CPU (fml/cpu.hh), GPU (fml/gpu.hh), and MPI (fml/mpi.hh) backends. These include all relevant sub-headers. These are "frozen" in the sense that they will not move and will always include everything. However, as more namespaces are added, those too will be included in the super headers. The headers one folder level deep (e.g. those in fml/cpu) are similarly frozen, although more may be added over time. Headers two folder levels

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

API

  • Frozen: Existing APIs will not be developed further.
    • none
  • Stable: Existing APIs are not expected to change. Some new features may be added slowly.
    • cpumat/gpumat/mpimat classes
    • copy namespace functions
    • linalg namespace functions (all but parmat)
  • Stabilizing: Core class naming and construction/destruction is probably finalized. Function/method names and arguments are solidifying, but may change somewhat. New features are still being developed.
    • dimops namespace functions
  • Evolving: Function/method names and arguments are subject to change. New features are actively being developed.
    • stats namespace functions
  • Experimental: Nothing is remotely finalized.
    • parmat - all functions and methods

Internals are evolving and subject to change at basically any time. Notable changes will be mentioned in the changelog.

Philosophy and Similar Projects

Some similar C/C++ projects worth mentioning:

These are all great libraries which have stood the test of time. Armadillo in particular is worthy of a look, as it has a very nice interface and very extensive set of functions. However, to my knowledge, all of these focus exclusively on CPU computing. There are some extensions to Armadillo and Eigen for GPU computing. And for gemm-heavy codes, you can use nvblas to offload some work to the GPU, but this doesn't always achieve good performance. And none of the above include distributed computing, except for PETSc which focuses on sparse matrices.

There are probably many other C++ frameworks in this arena, but none to my knowledge have a similar scope to fml.

Probably the biggest influence on my thinking for this library is the pbdR package ecosystem for HPC with the R language, which I have worked on for many years now. Some obvious parallels are:

The basic philosophy of fml is:

  • Be relatively small and self-contained.
  • Follow general C++ conventions by default (like RAII and exceptions), but give the ability to break these for the sake of performance.
  • Changing a code from one object type to another should be very simple, ideally with no changes to the source (the internals will simply Do The Right Thing (tm)), with the exception of:
    • object creation
    • printing (e.g. printing on only one MPI rank)
  • Use a permissive open source license.

About

Fused Matrix Library

Topics

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages